Inspiration
The AI industry is obsessed with building increasingly complex "agentic" systems. However, I noticed a massive gap: developers often blindly implement heavy agents without realizing the severe latency and cost penalties.
I wanted to answer the ultimate question:
When is a complex agent actually worth the overhead compared to a raw LLM?
I was inspired to build Agentic Race, a visceral, split-screen battleground that visualizes this exact tradeoff in real time, making AI benchmarking educational, transparent, and fun.
What it does
Agentic Race pits two AI architectures against each other simultaneously:
- Baseline Agent: A raw, zero-shot LLM call optimized purely for speed.
- Structured Agent: A complex workflow that dynamically plans, searches the web for real-time facts, reflects on its own answers, and refines its output.
Users enter a prompt, and a live race begins. As both agents stream their responses to the UI, a live stopwatch and token counter tick up in real time.
Once both agents finish, a third independent Judge AI evaluates them on:
- Accuracy
- Completeness
- Helpfulness
Finally, the app calculates a dynamic Cost Efficiency Score, providing a quantitative way to evaluate whether the Structured Agent's additional reasoning was worth its latency and token overhead.
How I built it
I built the backend in Python with FastAPI to handle asynchronous Server-Sent Events (SSE), streaming tokens instantly to the client.
The Structured Agent is powered by LangGraph, utilizing a multi-node state machine that integrates Tavily for live web searches and temporal awareness, allowing the agent to understand the current date and verify whether events have already occurred.
The frontend is built with Next.js 16 and React, utilizing Zustand for complex global state management to handle concurrent live streams without performance drops.
To power the AI, I utilized Groq for lightning-fast inference.
Challenges I ran into
My biggest challenge was maintaining stability at high speeds. When the Structured Agent performed heavy web searching, planning, and reflection within a few seconds, it completely exceeded the standard Tokens-Per-Minute (TPM) rate limits of the LLMs.
I also encountered severe JSON parsing bugs. When forcing reasoning models to output strict JSON for the Judge and Planning modules, they would occasionally generate invisible <think> tags or run out of context window space, resulting in 413 Request Entity Too Large errors and empty responses.
To solve this, I architected a Distributed Model System. I split the API load across three distinct models:
openai/gpt-oss-120bfor the Baseline's raw speed.qwen/qwen3.6-27bfor the Structured Agent's complex internal reasoning.qwen/qwen3.8-27bfor the Judge AI, selected for more reliable structured JSON generation.
By distributing the load across different API rate limit buckets and aggressively clamping web search context payloads, I was able to significantly reduce these bottlenecks.
Accomplishments I'm proud of
The Cost Efficiency Score
I designed a custom mathematical formula to measure the tradeoff between answer quality, latency, and token usage:
Cost Efficiency = Accuracy² / (Latency × log₁₀(Tokens))
This gives the benchmark a quantitative way to determine whether the additional complexity of an agentic workflow actually provides enough value to justify its overhead.
The "View Thinking" Modal
I built a highly transparent UI that allows users to open a modal and watch the agent's reasoning trace, web searches, and self-reflection loops as they happen.
The UX/UI
The cyberpunk and retro arcade aesthetic is highly responsive, with layout shifting, pulsing neon indicators, and micro-animations that make AI benchmarking feel more like a video game.
What I learned
I learned that agents are not a silver bullet.
For creative writing or standard logic tasks, the Baseline Agent can win easily on cost and latency. However, for tasks requiring real-time facts, the Structured Agent becomes significantly more valuable for reducing hallucinations.
I also learned deep lessons in productionizing LLM applications, particularly around safely enforcing structured JSON outputs, managing context window limitations, handling rate limits, and designing reliable streaming architectures.
What's next for Agentic Race
I want to expand the arena further. My next steps include:
- Adding new agent architectures, such as multi-agent debates and memory-enabled agents.
- Integrating more LLM providers, including OpenAI and Anthropic, so users can race different models and providers against each other.
- Building a global leaderboard for community-submitted edge-case prompts.
- Expanding the benchmark metrics to provide deeper insights into the quality, latency, and cost tradeoffs of different AI architectures.
Log in or sign up for Devpost to join the conversation.