Inspiration
DDoS attacks are one of the simplest ways to take a service down, and that's exactly what makes them so dangerous. Flood a target with enough packets, whether that's UDP, SYN, or ICMP, and the server's CPU, memory, or network stack eventually just gives up before it gets a real chance to defend itself. We found that even a modest flood of 400K packets per second is enough to bring a standard kernel-based Linux system to its knees, often before a single firewall rule even has time to fire.
What bothered us was how broken the existing options are for anyone who isn't a massive company. Cloud scrubbing services like Cloudflare or AWS Shield Advanced work fine, but they cost anywhere from $3K to $15K a month, route your traffic through a third party, and add latency along with a dependency you don't really control. Dedicated hardware appliances from companies like F5 or Radware solve the problem too, but they come with a $100K+ price tag upfront, which is massive overkill for a mid-size operator. Even iptables or nftables, the "free" route most people fall back on, is still kernel-space filtering. The overhead saturates at around 200K pps, so the CPU is still very much in the line of fire.
That got us asking a simple question: what if you could stop malicious packets before they ever reached the kernel in the first place?
What it does
PacketForge is a DPDK-based kernel-bypass DDoS scrubber. It intercepts traffic right at the NIC, before the Linux kernel ever gets to process a single packet.
Incoming traffic gets classified against configurable Lua rules in userspace. If a packet matches a blocked CIDR range or breaks a rate limit, it's dropped instantly and at hardware speed, without ever touching the protected server's CPU. Clean traffic is then forwarded out through a separate egress DPDK port directly to the protected server's MAC address.
Here's what we measured on commodity QEMU/KVM hardware:
- 470,000+ pps sustained throughput
- 100% drop rate once a block rule is active
- 0 ms to activate a new block rule (it's atomic)
- 0% CPU impact on the protected server
- The protected API stayed fully responsive the entire time, even during the simulated flood
- Zero kernel involvement throughout the attack
Because it runs on regular commodity hardware instead of proprietary appliances or a monthly cloud subscription, PacketForge brings scrubber-grade protection within reach of operators who would otherwise be stuck choosing between a service they can't afford and a kernel firewall that falls over under real load.
How we built it
The core of the pipeline runs on DPDK (Data Plane Development Kit), which lets us bind NICs directly to userspace through vfio-pci and skip the kernel's networking stack entirely. We use two NICs, one dedicated ingress port for inspecting incoming traffic and one egress port for forwarding clean packets to the protected server.
Packet classification happens through Lua-defined rules. We went with Lua specifically because it lets block rules, like CIDR ranges and rate limits, get added or changed atomically at runtime. That means zero downtime, no reload needed, and no dropped packets while a rule update happens.
To test all of this without needing actual bare-metal hardware, we built and ran the entire pipeline inside a QEMU/KVM virtual lab. One VM ran PacketForge with two DPDK-bound NICs, and a second VM played the role of both the attacker and the protected server.
Challenges we ran into
Running a DPDK scrubber inside a virtualized lab threw up three problems that you won't find answers to in any standard DPDK documentation, mostly because those docs assume you're working on bare-metal.
1. No IOMMU in QEMU/KVM. DPDK normally needs hardware IOMMU support to safely bind NICs to userspace through vfio-pci. The IOMMU is what guarantees a userspace driver can only access its own memory region. On a real server this just works out of the box, but QEMU/KVM doesn't have a real IOMMU, so the bind either failed silently or threw an IOMMU not found error. We got around this by enabling unsafe_noiommu_mode on the vfio kernel module. It bypasses the IOMMU requirement, and yes, it costs you the memory isolation guarantees, but that tradeoff is perfectly fine in a controlled VM lab. It's just not something you'd want to run in production without an actual IOMMU.
2. libvirt's nwfilter was quietly eating our test traffic. To simulate a realistic flood, we first tried hping3 --rand-source to spoof the source IPs. The PacketForge dashboard showed 0 pps the entire time, even though the flood was clearly running on our end. After a long and frustrating debugging session, we figured out that libvirt's built-in nwfilter was silently dropping every spoofed packet at the hypervisor bridge. It enforces a rule that an outgoing packet's source IP has to match the IP assigned to that VM's interface, so our attack traffic never even made it to the DPDK ingress port. The fix was simple once we found it: drop the IP spoofing and generate the flood from the VM's real interface IP instead. It's still a legitimate high-rate flood, and PacketForge still caught and blocked the source CIDR correctly.
3. ARP and MAC resolution for the egress forwarding path. Since DPDK skips the kernel entirely, PacketForge has no access to the kernel's ARP table, which means it can't dynamically figure out where to forward clean packets. Get the destination MAC wrong, and packets either go to the wrong place or just get broadcast. We solved this by hardcoding the protected server's MAC address directly into config.yaml. There was also a reverse version of this problem: the protected server's VM couldn't ARP-resolve PacketForge's DPDK-owned ingress IP, since a NIC owned by DPDK can't respond to kernel-level ARP requests at all. We fixed that by adding a static ARP entry on the protected server's VM that points PacketForge's ingress IP straight to its DPDK NIC's MAC address.
Accomplishments that we're proud of
- Getting a full DPDK kernel-bypass pipeline running end to end inside a virtualized lab, despite most DPDK tooling assuming you have bare-metal hardware
- Sustaining 470,000+ pps with a 100% drop rate on malicious traffic while measuring 0% CPU impact on the protected server
- Pulling off atomic, zero-downtime rule updates, where a new block rule takes effect immediately with no packets lost during the update
- Tracking down and fixing three separate, poorly documented virtualization quirks (IOMMU, libvirt's nwfilter, and DPDK/ARP) that had no clear answers anywhere in the standard DPDK docs
What we learned
Working at the packet level showed us just how much invisible infrastructure sits between an attacker and a target. Hypervisor bridge filters, kernel ARP tables, IOMMU memory isolation, any one of these layers can silently break your entire test setup without throwing a single error. We walked away with a much deeper understanding of the tradeoffs between kernel-space and kernel-bypass networking, and why真正 wire-speed packet processing means rethinking things like ARP resolution or CPU involvement that are normally just handled for you behind the scenes.
What's next for PacketForge
- Validating performance on real bare-metal hardware with an actual IOMMU, instead of relying on
unsafe_noiommu_mode - Expanding the Lua rule engine to support behavioral and anomaly-based detection, not just CIDR ranges and rate limits
- Adding multi-NIC and multi-server support so we can protect more than one downstream server at a time
- Building a management dashboard or API so block rules can be added without editing config files directly
Log in or sign up for Devpost to join the conversation.