Inspiration
Municipal transit systems generate massive streams of real-time telemetry, often captured via the GTFS-Realtime (GTFS-RT) format. Traditionally, parsing and processing these streams has relied on high-level, garbage-collected languages. While these languages offer convenience, they can introduce unpredictable latencies and garbage collection pauses that complicate high-frequency tracking, real-time simulation, and low-latency IPC distribution.
The inspiration behind VelocityRT was to explore the limits of safe systems programming. We set out to design a processing engine capable of digesting high-frequency telemetry streams with predictable, nanosecond-scale latencies. The goal was to prove that a system could achieve high-performance characteristics—such as zero dynamic memory allocations in hot paths and constant-time $O(1)$ parsing—while maintaining absolute memory safety under Rust’s #![forbid(unsafe_code)] directive.
What it does
VelocityRT is a high-assurance, low-latency GTFS-Realtime telemetry parsing and stream-processing engine. It is designed to ingest raw binary telemetry feeds and process them with minimal overhead.
The core system components perform several key functions:
- Zero-Copy Parsing: Converts raw binary stream bytes directly into structured telemetry data without heap allocations.
- SIMD Evaluation: Uses vectorized operations to scan and evaluate batches of vehicle telemetry (such as detecting speeding vehicles) at high throughput.
- High-Assurance Cryptography & Security: Signs and verifies telemetry batches using a BLAKE3 keyed MAC, featuring constant-time signature verification to mitigate timing side-channel attacks and automatic memory scrubbing for sensitive keys.
- Low-Latency Storage & IPC: Distributes processed telemetry over shared memory IPC using zero-copy loans and persists data via a hybrid storage engine combining an in-memory cache with a log-structured merge-tree (LSM) write-ahead log.
- Visual Diagnostics: Includes an optional background fetcher and a web dashboard to visualize active vehicles moving in real time.
How we built it
VelocityRT is written in safe Rust. We utilized a highly modular architecture based on the Facade & Adapter Pattern to isolate external dependencies from the core domain logic:
- Facade & Core Layer (
zc_gtfs_rt_core): This layer exposes clean, stable public abstractions (likeTelemetryEngineandVehicleTelemetry). External users interact with these clean APIs without directly touching low-level libraries. - Adapter Layer (
src/internal/): Third-party libraries are encapsulated behind internal Rust traits (such asStorageAdapter,IpcAdapter, andCryptoAdapter). This decouples our core logic from upstream changes in libraries likefjall,iceoryx2,zstd, orblake3. - Data Layout & Alignment: To prevent CPU false sharing across multi-threaded worker pipelines, all core telemetry structures are aligned to 64-byte cache boundaries using
#[repr(C, align(64))]. Zero-copy parsing is achieved by safely casting byte slices using thebytemuckcrate. - Instruction-Level Parallelism (ILP): The batch processing logic utilizes unrolled loop structures to assist the compiler in generating efficient SIMD instructions for evaluating large arrays of telemetry data.
Challenges we ran into
Building a system designed for high performance under strict safety constraints presented several notable technical challenges:
- Strict Memory Alignment for Zero-Copy Casting: Directly casting raw byte slices into Rust structures without allocating memory requires exact alignment. Ensuring that our structs met
bytemuck's safety invariants—without relying onunsafeblocks—required careful structural layout design, explicit padding, and rigorous compile-time checks. - Mitigating CPU False Sharing: During high-frequency parallel processing, multiple CPU cores updating adjacent memory addresses can trigger cache line bouncing, significantly degrading throughput. We addressed this by enforcing a 64-byte alignment on telemetry structures, matching typical CPU cache lines.
- Decoupling Fast-Evolving Dependencies: Integrating high-performance crates like
iceoryx2(for shared-memory IPC) andfjall(for storage) introduced the risk of breaking API changes. Designing the internal adapter layers to wrap these dependencies cleanly without leaking their complex types into our public API surface required multiple architectural refactorings.
Accomplishments that we're proud of
- Nanosecond-Scale Processing: We achieved an average parsing latency of approximately $1.84 \text{ ns}$ per batch in micro-benchmarks by relying entirely on zero-copy casting.
- Efficient Batch Evaluation: The 8-way loop unrolling approach allows the telemetry dispatcher to evaluate batches at a rate of over $2 \times 10^9$ items per second on modern hardware.
- High-Assurance Execution: The entire codebase operates under a
#![forbid(unsafe_code)]restriction, demonstrating that ultra-low-latency processing does not require bypassing Rust's safety guarantees. - Side-Channel Protections: We successfully integrated cryptographic verification that runs in constant-time ($1.08 \ \mu\text{s}$ per batch) alongside automatic zeroization of cryptographic keys when they are dropped from memory.
What we learned
- The Impact of Cache Layouts: We learned that CPU cache line alignment (
align(64)) and data locality often influence performance more than micro-optimizing individual instruction blocks in modern multi-threaded architectures. - API Isolation Discipline: Implementing the Facade and Adapter patterns early saved significant development effort. Isolating complex third-party APIs behind simple, domain-specific traits kept the compiler's build graph clean and simplified writing tests.
- Safe Zero-Copy Abstractions: We discovered that with careful struct design, crates like
bytemuckallow developers to achieve C-style casting performance safely, showing thatunsafeis rarely mandatory for high-performance parsing.
What's next for VelocityRT (zc_gtfs_rt_parser)
- AVX-512 and ARM Neon Target Optimizations: We plan to explore target-specific vectorization parameters to further improve SIMD processing efficiency on diverse CPU architectures.
- Multi-Feed Consolidation: Expanding the processing pipeline to ingest, merge, and reconcile multiple heterogeneous GTFS-RT feeds concurrently.
- Real-World Network Testing: Deploying the binary daemons against live, erratic municipal transit feeds to evaluate how the system handles packet loss, malformed packets, and network-induced jitter under real-world operating conditions.
Built With
- aes-gcm
- argon2
- axum
- blake3
- bytemuck
- criterion
- fjall
- gtfs-rt
- hmac
- iceoryx2
- plotters
- quick-cache
- rayon
- reqwest
- rust
- serde
- sha2
- shared-memory-ipc
- simd
- subtle
- tokio
- zero-allocation
- zeroize
- zstd
Log in or sign up for Devpost to join the conversation.