Intelligent AI Router
Modern AI applications often face a difficult trade-off: use a small model and get low latency and low cost, or use a powerful model and get better reasoning at a much higher computational and financial cost.
We wanted to explore whether an AI system could make that decision dynamically for every request instead of sending everything to the most expensive model.
That idea became Intelligent AI Router — a multi-tier AI inference system designed and optimized for AWS Graviton Arm CPU infrastructure.
The Core Idea
The core idea is simple:
Not every question deserves the same amount of compute.
Our system uses a local fine-tuned Qwen-based classifier to estimate the complexity of an incoming request and dynamically routes it through progressively more capable models:
flowchart TD
A[User Request] --> B[Local Qwen Router]
B --> C[v1 Gemma 3 1B]
B --> D[v2 Qwen3-4B]
B --> E[v3 GPT-5.6 Luna]
Simple requests can be handled by the lightweight local Gemma 3 1B model, moderately complex requests can use the local Qwen3-4B, and difficult requests can escalate to GPT-5.6 Luna through OpenRouter.
The goal is to minimize unnecessary cloud inference while preserving answer quality.
What Inspired Us
The inspiration came from a practical observation while working with AI systems: model capability is often treated as a binary choice when it should really be a routing problem.
A user asking:
"What is 17 × 8?"
doesn't need the same computational resources as someone asking:
"Design a fault-tolerant payment architecture for 10,000 transactions per second and compare the consistency, ordering, idempotency, recovery, latency, and cost trade-offs."
Sending both requests to a large reasoning model is wasteful.
We therefore asked:
Can an AI system determine how much intelligence a request actually needs?
This led us to design a cascading architecture where computation increases only when necessary.
The Arm/Graviton environment made the problem even more interesting. Instead of treating CPU inference as simply a fallback, we wanted to demonstrate that small, optimized local models can perform useful inference directly on Arm infrastructure, reducing dependence on cloud inference for appropriate workloads.
How We Built It
The system was built as a containerized multi-service application running on an AWS Graviton c8g.xlarge instance.
The architecture consists of several key components.
1. Local Complexity Router
A fine-tuned Qwen3-based classifier runs locally on Arm CPU.
Its job is not to answer the user's question. Instead, it evaluates the request and determines which generation tier is appropriate.
The classifier produces information such as:
- complexity score
- confidence
- task type
- routing decision
- reasoning for the decision
This separates "How difficult is this?" from "What is the answer?"
2. v1 — Gemma 3 1B
The first generation tier is a locally hosted fine-tuned Gemma 3 1B model.
It is optimized for simple, direct-answer requests.
For example:
"What is 2 + 2?"
was handled locally in approximately 158 ms, with around 57 tokens/sec in our deployment test.
This demonstrates the main economic advantage of the architecture: extremely simple tasks can remain entirely on local Arm infrastructure.
3. v2 — Qwen3-4B
We introduced Qwen3-4B as an intermediate generation tier.
This became important because some questions are clearly more complex than what we want to give to a 1B model, but don't necessarily require a large cloud model.
Qwen3-4B runs locally using llama.cpp and a Q4_K_M GGUF checkpoint.
One of the interesting challenges we encountered was that Qwen3 naturally enabled its reasoning behavior. When tested normally, it could spend its entire output budget generating internal reasoning while returning an empty final answer.
We solved this by explicitly disabling thinking:
chat_template_kwargs = {
"enable_thinking": False
}
This turned Qwen3-4B into a much more appropriate fast-response middle tier.
4. v3 — GPT-5.6 Luna
The final tier is our strongest model, accessed through OpenRouter.
It acts as the quality safety net.
If a request is too difficult for the local tiers, or a local response fails validation, the system can escalate rather than returning a poor answer.
This gives us a fundamental design principle:
Local models optimize cost and latency; the cloud model protects quality.
Escalation Architecture
The final routing pipeline works as a cascade:
flowchart TD
A[User Request] --> B[Qwen Classifier]
B --> C[Complexity Score]
C -->|Simple| D[Gemma 1B]
C -->|Complex| E[Qwen 4B]
D --> F{Verify output}
E --> G{Verify output}
F -->|Pass| H[Return]
F -->|Reject| E
G -->|Pass| I[Return]
G -->|Reject| J[GPT 5.6- Luna]
E --> G
This creates a graceful degradation path instead of relying on a single model.
What We Learned
One of the biggest lessons was that model selection is only part of the problem.
The inference infrastructure itself matters enormously.
We learned how to:
- deploy multiple AI models on Arm CPU infrastructure
- serve GGUF models through llama.cpp
- build ARM64 Docker images
- connect independently managed containers through Docker networks
- expose OpenAI-compatible model APIs
- implement dynamic model routing
- measure latency and token throughput
- implement response verification and escalation
- maintain request-level observability
- track estimated inference cost and savings
- distinguish model failures from infrastructure failures
We also learned that a model's theoretical capabilities don't necessarily translate directly into application performance.
For example, our initial Qwen3-4B test appeared to fail because the model returned:
finish_reason: length
content: ""
while spending the token budget inside reasoning_content.
The model itself was healthy. Our serving configuration was wrong for the intended use case.
That was an important engineering lesson: inference optimization is not just about choosing a checkpoint. Prompt format, chat templates, context size, token limits, threading, and runtime configuration all affect the final application.
Challenges We Faced
ARM64 Compatibility
Running the entire stack on AWS Graviton introduced architecture-specific considerations.
We had to ensure that the Docker images, inference runtime, Python dependencies, and compiled components were compatible with ARM64.
Limited CPU Resources
Unlike GPU inference environments, our local models were running on CPU.
This made threading, context size, parallelism, quantization, and model selection important performance considerations.
Docker Networking
At one point the API could not resolve the Qwen model because the model container and API were on different Docker networks.
We diagnosed the problem by inspecting Docker network assignments and eventually connected the services correctly.
Incorrect Healthcheck
Our manually created Qwen container initially appeared unhealthy even though the model was working.
The reason was straightforward:
- Healthcheck → port
8080 - Qwen server → port
8081
The healthcheck was testing the wrong port.
Correcting the healthcheck made the container report its actual state.
Reasoning Token Consumption
Qwen3-4B initially consumed its token budget on reasoning instead of producing a final answer.
Disabling thinking through the chat template configuration fixed the issue.
Disk Constraints
Building the ARM64 stack also exposed a practical infrastructure problem: Docker build cache and model layers consumed significant disk space.
At one point the instance reached approximately 92% disk utilization, causing a Docker build to fail with:
no space left on device
We recovered by cleaning unused Docker build cache and then rebuilt only the API instead of unnecessarily rebuilding the large model images.
Results
We successfully deployed and tested the complete local routing stack.
Our observed results included:
| Tier | Model | Example Performance |
|---|---|---|
| v1 | Gemma 3 1B | ~158 ms |
| v2 | Qwen3-4B | ~7.9 s for a 128-token response |
| v3 | GPT-5.6 Luna | Cloud fallback |
We also verified that requests were being distributed between the local tiers rather than everything being sent to the cloud.
One deployment metrics snapshot showed:
| Metric | Value |
|---|---|
| Total requests | 4 |
| v1 requests | 2 |
| v2 requests | 2 |
| v3 requests | 0 |
| Local handling rate | 100% |
| Escalated requests | 0 |
| Error rate | 0% |
This demonstrated that the router was successfully keeping suitable workloads on the local Arm infrastructure.
The Bigger Idea
The most important takeaway from this project is that AI optimization doesn't necessarily mean making one model faster.
It can mean deciding when not to use the expensive model at all.
If the probability that a lightweight model can successfully answer a request is high, then the system should use it.
If confidence falls or the response fails validation, computation can increase progressively.
Conceptually, the system is optimizing something like:
$$ \min \mathbb{E}[\text{latency}] + \lambda \mathbb{E}[\text{cost}] $$
subject to:
$$ P(\text{acceptable answer}) \geq Q $$
where the router dynamically chooses the smallest model capable of satisfying the required quality.
That is the central idea behind our project:
Use the minimum amount of intelligence necessary — and escalate only when the problem demands more.
By combining fine-tuned local models, Arm CPU inference, dynamic routing, verification, and intelligent escalation, we built a system that treats AI inference as an optimization problem rather than simply a model-selection problem.
Built With
- amazon-web-services
- arm64
- docker
- face
- fastapi
- gemma
- graviton
- hugging
- openrouter
- peft
- python
- pytorch
- qwen3
Log in or sign up for Devpost to join the conversation.