I opened DistroWatch after a long time and came across HardenedBSD. Comparing it with FreeBSD and OpenBSD, which I’m more familiar with, led me to explore the hardening techniques used by HardenedBSD and OpenBSD. I decided to document a few of those techniques through small experiments, following each defense from the build option that enables it to the change it makes in program behavior.
I’m using C because a small program lets me trace a memory operation into the generated instructions, making both an unsafe copy and the checks added to defend it visible. C is central to the Linux kernel, SQLite, Redis, NGINX, and the Asterisk telephony platform; it is also the native extension language for CPython. Understanding these mechanisms is therefore useful even when the application itself is written in another language.
The environment
I’m going to use Ubuntu 26.04 LTS on 64-bit Arm Linux, running in Docker on my M4 Pro MacBook, and compile the examples for the Armv8-A instruction baseline. Arm calls this 64-bit execution state AArch64, while Docker identifies the platform as linux/arm64.
Canonical maintains the official Ubuntu image, a common base for development and server containers. Alongside Ubuntu 26.04’s GitHub Actions and Azure Pipelines runner images, that makes it a practical choice for an experiment intended to transfer to build and deployment environments.
I’ll start in the source directory on my Mac:
docker pull --platform linux/arm64 ubuntu:26.04
docker image inspect ubuntu:26.04 \
--format '{{index .RepoDigests 0}}' | tee image-digest.txt
docker run --rm -it --platform linux/arm64 \
--mount "type=bind,src=$PWD,dst=/work" -w /work ubuntu:26.04 bash
Inside the container, I’ll install GCC 15 and the tools needed to build and inspect the examples:
apt-get update
apt-get upgrade -y
apt-get install -y --no-install-recommends gcc-15 libc6-dev binutils
CC=gcc-15
GNU binutils supplies the linker, which combines compiled code into an executable, and the inspection tools used below. Since the GNU C Library, glibc, supplies calls such as memcpy, I’ll record its version alongside those of the compiler and linker:
{
"$CC" -dumpfullversion
"$CC" -dumpmachine
getconf GNU_LIBC_VERSION
ld --version | head -n 1
dpkg-query -W -f='${binary:Package}\t${Version}\n' gcc-15 libc6 binutils
} | tee toolchain.txt
uname -m
uname -r
The current GCC, glibc, and binutils package revisions imply the output below; it is an expected example, not output captured from a Docker run:
15.2.0
aarch64-linux-gnu
glibc 2.43
GNU ld (GNU Binutils for Ubuntu) 2.46
binutils 2.46-3ubuntu2
gcc-15 15.2.0-16ubuntu1
libc6:arm64 2.43-2ubuntu2.4
I’ll retain the actual output alongside the image digest to record the build environment. uname -m should report aarch64, while uname -r reports Docker Desktop’s Linux VM kernel, since the Ubuntu image supplies the userspace rather than the kernel.
I’ll keep the commands in one Bash session, without set -e, so a deliberately failing probe doesn’t prevent the next one from running. The editable Compiler Explorer examples use AArch64 GCC 15.2 with separate headers and libraries, so their assembly is a comparison rather than a reproduction of Ubuntu’s build. The Docker runtime outcomes below remain expectations to verify.
The program and the baseline
I’m going to use a small command-line program that takes one text argument, copies it into a fixed-size record field, and prints the record. It models the point where a parser or record formatter moves variable-length input into storage with a fixed capacity, while keeping the code small enough to inspect.
The format uses a 32-byte field with room for up to 31 bytes of text and a terminating zero. Initializing the field to zero means a shorter copy already has that terminator, but I’m deliberately omitting the length check from the first build to see how the defenses respond when the input doesn’t fit. I’ll save this as record.c:
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
__attribute__((noinline))
static int show_record(const char *input)
{
char field[32] = {0};
size_t length = strlen(input);
#ifdef CHECK_LENGTH
if (length >= sizeof field) {
fputs("record too long\n", stderr);
return 2;
}
#endif
memcpy(field, input, length);
printf("record: %.32s\n", field);
return 0;
}
int main(int argc, char **argv)
{
if (argc != 2) {
fputs("usage: record TEXT\n", stderr);
return 2;
}
#ifdef PRINT_ADDRESS
if (strcmp(argv[1], "--address") == 0) {
printf("show_record: %p\n", (void *)show_record);
return 0;
}
#endif
return show_record(argv[1]);
}
CHECK_LENGTH enables the repair, while the initial build leaves memcpy trusting the supplied length. I’ve bounded the print to 32 bytes so a second bug that reads beyond the field doesn’t obscure the write being examined. PRINT_ADDRESS enables a diagnostic for the later address comparison.
The local array is stored on the stack, which also holds saved state for function calls. I’ve used noinline to keep show_record separate for inspection with objdump, which disassembles the binary into assembly text.
To establish a baseline, I’ll explicitly disable the defenses I’m about to examine; otherwise, compiler and distribution defaults could obscure which option changes the behavior. A Bash helper generates inputs of an exact byte length using built-in printf and string substitution:
ulimit -c 0
base=(
-std=c17 -O2 -g -Wall -Wextra -march=armv8-a -DPRINT_ADDRESS
-U_FORTIFY_SOURCE # unchecked library copies
-fno-stack-protector # no stack guard
-mbranch-protection=none # no Arm hardware branch checks
-fno-pie -no-pie # fixed executable addresses
-Wl,-z,norelro,-z,lazy,-z,execstack # writable linker data; executable stack
)
repeat_x() {
local padding
printf -v padding '%*s' "$1" ''
printf '%s' "${padding// /x}"
}
long_input=$(repeat_x 80)
short_overflow=$(repeat_x 33)
"$CC" "${base[@]}" record.c -o baseline
./baseline hello
./baseline "$long_input"
The valid input prints record: hello, while the 80-byte input writes beyond the field and may corrupt surrounding stack state or crash. Because an out-of-bounds write is undefined behavior, C specifies no result for it: a normal exit wouldn’t establish safety, just as a crash wouldn’t establish that an attacker can control execution.
Inspect the baseline on Compiler Explorer.
The next five experiments add defenses one at a time, pairing each with a probe of the risk it addresses. I’ll distinguish between detecting corruption after a write, rejecting the write before it occurs, and limiting how corrupted state can be used. Hardware-assisted protections will be the subject of a future post.
1. Stack canaries
Writing beyond a stack buffer can corrupt nearby local data or saved call state, including a return address that tells the processor where to resume after a function finishes. A stack smashing attack tries to use that overwrite to redirect execution by changing such state. The unsafe copy establishes the write; whether it can provide that control depends on the generated layout and other protections.
A stack canary is a guard value saved in a function’s stack storage and checked before returning to detect an overwrite that reaches it. I’ll enable it with -fstack-protector-strong:
canary_flags=(-fstack-protector-strong)
"$CC" "${base[@]}" "${canary_flags[@]}" record.c -o canary
./canary "$short_overflow"
objdump -d --disassemble=show_record canary
In the linked GCC 15.2 example, field occupies offsets 24–55 from the stack pointer, which tracks the current stack position, and the guard begins at offset 56. The 33rd byte therefore reaches the guard, leading me to expect *** stack smashing detected *** followed by an abort signal, SIGABRT. I’ll inspect the Ubuntu build before assuming that it has the same layout.
The linked instructions load __stack_chk_guard, save a copy, and call __stack_chk_fail if the value no longer matches. Because the comparison runs after both memcpy and printf, corruption that misses the guard or affects data used before the check can escape this defense.
Compare baseline and canary assembly.
2. Fortification
A canary checks for damage after the copy, so it cannot protect data used before that check. When the compiler supplies the destination’s capacity, glibc’s fortification can reject an oversized library copy before it writes. I’ll enable it with _FORTIFY_SOURCE=3 while retaining the canary:
fortify_flags=("${canary_flags[@]}" -D_FORTIFY_SOURCE=3)
"$CC" "${base[@]}" "${fortify_flags[@]}" record.c -o fortified
./fortified "$short_overflow"
objdump -d --disassemble=show_record fortified
The linked example calls __memcpy_chk with 32 in register x3, passing the destination capacity as the fourth argument. Because this check rejects the 33-byte length before copying, I expect *** buffer overflow detected *** and SIGABRT without reaching the canary check.
At level 3, the compiler can compute object sizes at runtime, extending the check to cases where the bound isn’t a compile-time constant. A library call can still remain unchecked when no usable bound is available, and ordinary pointer writes remain outside this protection.
Compare canary and fortified assembly.
3. Randomizing executable addresses
An attacker who can overwrite a return address still needs somewhere to redirect execution, and predictable code addresses make it easier to choose a target. Linux uses address space layout randomization (ASLR) to vary process addresses, but for the main executable to participate, it must support loading at different addresses as a position-independent executable (PIE).
I’ll compare addresses across fresh processes and use readelf, which inspects metadata in Linux’s Executable and Linkable Format (ELF) binaries, to check how the executable was built.
pie_flags=("${fortify_flags[@]}" -fPIE -pie)
"$CC" "${base[@]}" "${pie_flags[@]}" record.c -o pie
cat /proc/sys/kernel/randomize_va_space
for i in 1 2 3; do ./fortified --address; done
for i in 1 2 3; do ./pie --address; done
readelf -hW pie
With ASLR enabled, I expect a stable address from fortified and varying addresses from pie, whose ELF type DYN identifies the relocatable form used here. The build options permit relocation, while Linux chooses where to load the executable.
The three runs can show address variation, but they won’t measure how difficult an address is to guess. A leaked address can reveal the location despite randomization, which is why I’ll omit PRINT_ADDRESS from the final build.
4. A non-executable stack
An attacker who can both place instruction bytes on the stack and redirect execution to them can run those bytes if the stack is executable. I’ll remove execute permission to prevent that use of stack memory:
nx_flags=("${pie_flags[@]}" -Wl,-z,noexecstack)
"$CC" "${base[@]}" "${nx_flags[@]}" record.c -o nx
readelf -lW nx | awk '/GNU_STACK/'
The GNU_STACK metadata now requests RW, read and write permissions, instead of RWE, which also permits execution. Since the record overflow alone doesn’t demonstrate execution from the stack, I’ll use a separate probe that attempts exactly that, saved as stack.c:
#include <stdio.h>
int main(void)
{
volatile unsigned char code[] __attribute__((aligned(4))) =
{0xc0, 0x03, 0x5f, 0xd6}; /* A64: RET X30 */
__builtin___clear_cache((char *)code, (char *)code + sizeof code);
((void (*)(void))code)();
puts("stack code returned");
}
The probe calls a return instruction stored on the stack, using GCC/Linux behavior beyond portable C; RET X30 returns to the address saved in register x30. I’ve aligned the bytes to four bytes, as required for A64 instructions, and used volatile to keep them materialized. The builtin handles cache synchronization so the newly written bytes are visible to instruction fetch.
"$CC" "${base[@]}" stack.c -o stack-exec
"$CC" "${base[@]}" -Wl,-z,noexecstack stack.c -o stack-nx
./stack-exec
./stack-nx
I expect stack-exec to print stack code returned and stack-nx to fail with a memory-access signal, SIGSEGV, although a stricter kernel policy may reject both. Removing execute permission prevents execution from this destination, while corruption can still alter data or redirect execution through code that is already executable.
5. Protecting library-call addresses
The loader connects an executable to the libraries it uses, with some calls reading a stored function address from the Global Offset Table (GOT). Changing a writable entry can turn an ordinary library call into a call to an address chosen by the attacker.
Relocation read-only (RELRO) protection lets the loader mark selected relocation data read-only, while full RELRO also resolves library-call addresses at startup so the corresponding GOT entries no longer need to be updated later. I’ll enable both linker options:
linked_flags=("${nx_flags[@]}" -Wl,-z,relro,-z,now)
"$CC" "${base[@]}" "${linked_flags[@]}" record.c -o linked
readelf -lW linked | awk '/GNU_RELRO/'
readelf -dW linked | awk '/BIND_NOW|FLAGS/'
To test the protection directly, I’ll use a program that calls puts, overwrites the address used for that call, and calls puts again. I’ll save it as relro.c:
#include <stdint.h>
#include <stdlib.h>
#include <stdio.h>
#include <unistd.h>
static int replacement(const char *text)
{
(void)text;
const char message[] = "redirected call\n";
return write(STDOUT_FILENO, message, sizeof message - 1) < 0;
}
int main(int argc, char **argv)
{
if (argc != 2)
return 2;
puts("original call");
fflush(stdout);
volatile uintptr_t *slot =
(uintptr_t *)strtoull(argv[1], NULL, 0);
*slot = (uintptr_t)replacement;
puts("original call");
return 0;
}
This probe models a write to a chosen address, a capability the record overflow has not been shown to provide. I’ll compare partial RELRO, which leaves library-call entries writable for later resolution, with full RELRO, keeping fixed executable addresses so readelf’s relocation addresses match their runtime locations:
"$CC" "${base[@]}" -Wl,-z,relro,-z,lazy relro.c -o partial
"$CC" "${base[@]}" -Wl,-z,relro,-z,now relro.c -o full
for binary in partial full; do
slot=$(readelf -rW "$binary" |
awk '/JUMP_SLOT/ && /puts@/ {print "0x" $1; exit}')
test -n "$slot" || { printf 'puts slot missing\n'; break; }
./"$binary" "$slot"
done
I expect partial RELRO to print original call, then redirected call, while full RELRO should print the first line and fault when the probe tries to change the protected entry. This protects selected linker data; application function pointers, which store callable addresses, can remain writable elsewhere. The GNU linker reference specifies these options.
GCC and Clang beyond these defenses
GCC and Ubuntu’s Clang 22.1.2 both support these defenses, while their additional checks address different failure modes and have different build requirements.
Clang can check whether a call through a function pointer reaches a function with an allowed type, a form of control-flow integrity (CFI). A compatible type still doesn’t prove that the destination is the intended one. Most Clang CFI schemes require link-time optimization (LTO), which lets the compiler analyze code across source files while linking; shared-library boundaries need separate attention.
Clang’s SafeStack moves vulnerable stack objects away from return addresses and safely accessed local data, with support from additional runtime code. It doesn’t support building shared libraries with SafeStack instrumentation.
GCC offers checks on a function’s executed path, which address a different problem from checking the type of an indirect-call destination. Its stack scrubbing clears selected stack storage after use, with some modes changing how functions call one another.
GCC’s -fhardened bundles several production defenses, but I’m keeping explicit flags in this walkthrough so each change can be connected to its effect. The bundle can be inspected with gcc-15 --help=hardened.
Fixing the input check
The original input-handling error remains in every build above because the build options don’t enforce the record format. To enforce that rule, I’ll enable CHECK_LENGTH, retain the demonstrated defenses, and omit the address diagnostic:
"$CC" -std=c17 -O2 -g -Wall -Wextra -march=armv8-a \
-DCHECK_LENGTH -D_FORTIFY_SOURCE=3 -fstack-protector-strong \
-fPIE -pie -Wl,-z,relro,-z,now,-z,noexecstack \
record.c -o record
./record "$long_input"
./record "$(repeat_x 31)"
./record "$(repeat_x 32)"
./fortified "$(repeat_x 32)"
The repaired program accepts 31 bytes and rejects 32 or more before copying, reporting record too long with status 2. Fortification alone accepts 32 because the copy fits the destination, whereas the explicit length check leaves one of the zero-initialized bytes untouched to meet the record format.
Inspect the repaired build on Compiler Explorer.
Conclusion
Enabling hardening changes the executable’s behavior even when the faulty source remains the same: a checked library call can stop an oversized copy before it writes, while a stack guard can detect certain overwrites before a function returns. The compiler adds these checks to code whose author never wrote them explicitly, while linker settings and operating-system enforcement support address randomization and restrict writes or execution in selected memory regions.
The compiler can know that 32 bytes fit the destination without knowing that the format needs a terminating zero, which is why the input check remains necessary. Checks can also terminate the program, and corruption outside their coverage remains possible.
Applicable hardening options belong in the production build configuration, with their effects verified in the generated instructions and executable metadata. Source-level validation enforces the application’s rules; compiler, library, and linker protections provide additional barriers when a memory error survives review and testing.