What Is eBPF? Safe Programs Inside the Linux Kernel, Explained
eBPF lets you run small, verified programs inside the Linux kernel at hook points — syscalls, network packets, function entry — without kernel modules. How it works (verifier, JIT, maps, hooks), what it powers, bpftrace one-liners, and its limits and security implications.
eBPF lets you load small programs into the running Linux kernel and attach them to events — a system call, a packet arriving, a function being entered, a file being opened. The kernel verifies each program is safe before running it, so you get kernel-level visibility and control without writing kernel modules or rebooting.
It quietly powers a lot of modern infrastructure: container networking (Cilium), runtime security (Falco, Tetragon), continuous profilers, load balancers, and observability agents.
From packet filter to general-purpose engine
The original BPF (Berkeley Packet Filter, 1992) was a tiny virtual machine for filtering packets — it's what tcpdump 'port 443' compiles to. Extended BPF (eBPF), added from Linux 3.18 onwards, generalised it: more registers, 64-bit, maps for state, helper functions, and hooks all over the kernel. Today "BPF" usually means eBPF.
How it works
your program (C / Rust / bpftrace)
│ compile (clang/LLVM)
▼
eBPF bytecode ──bpf() syscall──▶ verifier ──▶ JIT ──▶ native code attached to a hook
│
maps ◀──────────┘ (shared with user space)
- Verifier. Before loading, the kernel statically checks the program: all memory access is bounds-checked, loops are bounded, it always terminates, it only calls allowed helpers, and it can't crash the kernel. Programs that don't pass are rejected. This is what makes eBPF safe in a way kernel modules aren't — and also what makes writing it fiddly.
- JIT compiler. Verified bytecode is compiled to native machine code for near-native speed.
- Hooks. The program attaches to an event source.
- Maps. Key–value data structures (hash maps, arrays, ring buffers, per-CPU arrays…) shared between eBPF programs and user space — how programs keep counters and send events out.
Where programs attach
| Hook | Fires on | Used for |
|---|---|---|
| kprobe / fentry | Entry/exit of (almost) any kernel function | Deep tracing, debugging |
| tracepoint | Stable, predefined kernel events | Reliable tracing across versions |
| uprobe / USDT | Functions in user-space programs | Tracing apps without changing them (e.g. TLS libraries, runtimes) |
| XDP | Packets at the network driver, before the kernel stack | DDoS filtering, very fast load balancing |
| TC | Packets in the traffic-control layer | Container networking, policy |
| cgroup hooks | Socket ops, device access per cgroup | Per-container network and device policy |
| LSM (BPF-LSM) | Security hooks | Custom mandatory access control |
| perf events | CPU sampling | Profiling (flame graphs) |
Try it: bpftrace one-liners
bpftrace is a high-level tracing language — awk for the kernel. On a test machine (needs root):
# Every program executed, with its arguments
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_execve { printf("%s -> %s\n", comm, str(args.filename)); }'
# Which processes open which files
sudo bpftrace -e 'tracepoint:syscalls:sys_enter_openat { printf("%s %s\n", comm, str(args.filename)); }'
# Histogram of read() sizes per process
sudo bpftrace -e 'tracepoint:syscalls:sys_exit_read /args.ret > 0/ { @[comm] = hist(args.ret); }'
# Count TCP retransmits by destination
sudo bpftrace -e 'tracepoint:tcp:tcp_retransmit_skb { @[ntop(args.daddr)] = count(); }'
The BCC tools collection includes ready-made tools like execsnoop, opensnoop, biolatency, tcpconnect and runqlat — the fastest way to answer "what is this server actually doing?"
What it powers
- Networking: Cilium implements Kubernetes networking, load balancing and network policy in eBPF instead of long iptables chains; XDP-based load balancers process millions of packets per second per core.
- Security: Falco and Tetragon watch syscalls and process behaviour in containers and alert on (or block) suspicious activity — a shell spawning in a web container, unexpected outbound connections. (Egress control for AI agents)
- Observability: continuous profilers and APM agents capture CPU profiles, latencies and even HTTP traffic with no code changes.
- Sandboxing building blocks: seccomp filters are classic BPF; BPF-LSM allows custom policies. (Container hardening)
Writing production eBPF
For anything beyond one-liners, the modern approach is CO-RE ("compile once, run everywhere") with BTF type information, so one compiled program adapts to different kernel versions. Libraries: libbpf (C), Aya (Rust), cilium/ebpf (Go). Expect to fight the verifier — keep programs small, push complex logic to user space via maps and ring buffers.
Limits and caveats
- Kernel version matters. Features arrived gradually; check what your production kernels support.
- Privileges. Loading most programs requires
CAP_BPF/CAP_SYS_ADMIN(plusCAP_PERFMON/CAP_NET_ADMINdepending on type). Unprivileged eBPF is disabled by default on most distributions because it has been a source of kernel vulnerabilities. - Security double edge. Anything that can load eBPF can observe (and in some hooks, modify) a lot — including other containers' traffic on the host. Treat eBPF capability like root. Containers should not get it.
- Overhead. Low, but not zero: a kprobe on a hot function fires millions of times a second.
- Not inside guests' kernels from the host. In VM-based isolation, the host's eBPF sees the VM as a process; tracing inside requires running eBPF in the guest kernel. (How KVM works)
The summary
- eBPF runs verified, JIT-compiled programs at kernel hooks without kernel modules.
- The verifier guarantees safety; maps share data with user space.
- Hooks span syscalls, kernel and user functions, packets (XDP/TC), cgroups and LSM.
- It powers Cilium, Falco, profilers and more;
bpftraceand BCC are the quickest way in. - Loading eBPF is a root-level privilege — never grant it to untrusted workloads.
EasySpawn isolates each server as its own VM with its own kernel, so low-level tooling inside one customer's workload can't observe another's. See how it works or join the waitlist.
Related: Linux Namespaces Explained · Container Networking Internals · Hardening Containers With seccomp and AppArmor · Finding Memory Leaks in Node.js
Keep reading
Linux Namespaces Explained: The Kernel Feature Containers Are Made Of
A container is a process with its own namespaces. What each of the eight Linux namespaces isolates — mount, PID, network, UTS, IPC, user, cgroup, time — how to build a container-like process by hand with unshare, inspect one with nsenter and /proc, and what namespaces do not protect against.
How KVM Virtualization Works: The Isolation Behind Cloud Servers
Most cloud VMs run on KVM. How hardware-assisted virtualization works — VT-x/AMD-V, the VMM, virtio, two-level page tables — what a VM boundary protects against compared with containers, the remaining risks (side channels, noisy neighbours, steal time), and what shared vs dedicated vCPUs mean.