High-Performance Packet Forensics: Leveraging GPU Co-Processors for Terabit-Scale Analysis
A high-level architecture for efficient packet trace analysis on GPU co-processors
The paper proposes a high-level GPU-based architecture for massively parallel packet classification and analysis. By utilizing a specialized Virtual Machine (VM) and a Domain-Specific Language (DSL), the system offloads computationally intensive filtering and visualization tasks to commodity GPU hardware, achieving terabit-scale classification speeds.
TL;DR
Network packet traces are getting massive, often reaching terabytes, making traditional CPU-based analysis tools painfully slow. This paper introduces a high-level architecture that uses Nvidia CUDA GPUs to accelerate packet classification and visualization. By treating the GPU as a specialized Virtual Machine (VM) and using a Domain-Specific Language (DSL), the authors achieve classification speeds that far exceed disk I/O capabilities, turning a compute-intensive task into a purely bandwidth-bound one.
The "Processor-Bound" Wall
In the world of cybersecurity and network research, packet traces (PCAP files) are the ultimate source of truth. However, as link speeds increase, these files grow to sizes that overwhelm standard tools.
The authors identify a critical flaw in current tools: Wireshark and similar software are "Processor-Bound." Even with fast SSDs, these tools cannot process packets quickly enough because the CPU handles every comparison and protocol parse sequentially. Prior GPU attempts like Gnort failed to solve this entirely because they still relied on the CPU for initial packet classification, creating a massive bottleneck.
Methodology: The GPU Virtual Machine Approach
The core innovation lies in shifting the entire classification logic into a GPU Classification Virtual Machine. Instead of hard-coding filters, the system uses a custom DSL (compiled via ANTLR) that generates optimized integer instructions.
1. High-Level Architecture
The system is split into three layers:
- User Layer (C#): Provides the DSL compiler and an OpenGL-based visualizer.
- Management Layer (C++): Handles heavy-duty I/O, buffering, and "File Mirroring" (reading from multiple drives simultaneously to maximize bandwidth).
- Accelerator Layer (CUDA): The VM that executes filter programs across thousands of GPU threads.

2. Handling Complexity and Divergence
One of the biggest challenges in GPGPU is Thread Divergence (where different threads follow different logic paths). The authors solved this by:
- Instruction Unification: Supplying all threads in a warp with the same instruction set.
- Register Caching: Using a local 64-byte packet cache to minimize expensive Global Memory fetches.
- Pass-based Scheduling: Automatically splitting complex protocols into multiple passes to avoid register spilling, keeping the execution on the fast "on-chip" memory.
Experimental Validation: Breaking the Bottleneck
The experiments highlight a paradigm shift. As shown in the performance comparison, the prototype application (CaptureFoundry) scales linearly with the speed of the storage medium, whereas Wireshark remains stagnant regardless of whether an HDD or SSD is used.

Key findings include:
- GPU Under-Utilization: The GPU is so fast that it sits idle for most of the time, waiting for the disk to provide data.
- Scalability: The architecture allows for increasing the complexity of filters (e.g., statistical analysis, data mining) without any drop in throughput, because the computation time is still lower than the I/O time.
Real-Time Visualization
To help humans make sense of billions of packets, the architecture generates Result Masks (optimized bit-arrays) and Index Files. This allows the UI to render traffic dynamics in real-time, effectively "skipping" to relevant data without re-parsing the whole terabyte-scale file.

Critical Insight & Future Outlook
The primary takeaway is that for massive data analysis, we should stop optimizing algorithms for the CPU and start building I/O-aware parallel architectures.
Limitations: The current model is optimized for "offline" batch processing. Applying this to live, low-latency streams (like a firewall) would require careful buffer management to avoid introducing significant lag.
Future Work: The authors envision a distributed network of these GPU-enabled sensors, providing a high-speed "eyes-on-the-network" capability across global infrastructures.
