Inspiration
Switzerland’s push for a sovereign, open-source AI with the Apertus models is a massive step forward for data privacy and local innovation. However, deploying massive 70B parameter models for every single user query is incredibly resource-intensive, expensive, and adds unnecessary latency for simple tasks. We were inspired to solve this by building a hybrid infrastructure layer. We asked: What if we could route simple queries to a local, on-device Apertus 8B model, and only escalate complex reasoning tasks to the cloud-hosted Apertus 70B model?
What it does
EdgeRouter is an intelligent AI inference gateway designed specifically for the Apertus ecosystem. It acts as a smart proxy between the user and the LLMs.
- Intelligent Routing Engine: Before any API call is made, EdgeRouter classifies the prompt’s complexity using an O(1) length gate and keyword classifier.
- Local Sovereignty: Simple, directive prompts (e.g., summarization, data extraction) are routed to a quantized Apertus 8B model running locally on the user's hardware. This ensures zero data leaves the machine, preserving absolute sovereignty and achieving near-zero network latency.
- Cloud Escalation: Complex reasoning tasks, coding requests, or advanced NLP workflows are automatically detected and routed to the massive Apertus 70B model hosted in the cloud.
- Live Telemetry Dashboard: The platform features a real-time analytics dashboard displaying Time-To-First-Token (TTFT), tokens/sec, and a visualizer showing the routing decision.
How we built it
We built EdgeRouter focusing on performance and seamless model swapping:
- The Router: The core routing logic is built in Node.js, utilizing a highly optimized O(1) local gate before any request is fired.
- Local Inference: We integrated local runtime support (via
llama.cppbindings) to run a quantized version of the Apertus 8B model directly on the host machine. - Cloud Inference: We connected to the cloud-hosted Apertus 70B API for escalated queries.
- The Dashboard: We built a beautiful 3-panel UI featuring an Agent Console, Live Benchmarks, and Session Analytics so developers can visually monitor the routing decisions and cost savings in real-time.
Challenges we ran into
The hardest challenge was designing a routing classifier that was faster than simply sending the request to the cloud. If our routing logic took 500ms to decide where to send the prompt, it would defeat the purpose of latency reduction. We solved this by avoiding LLM-as-a-judge approaches for the router, and instead building a blazing-fast, deterministic O(1) length and keyword heuristic gate. Another challenge was standardizing the prompt formatting between the 8B and 70B Apertus variants to ensure the user experienced no difference in output structure regardless of where the prompt was routed.
Accomplishments that we're proud of
We are incredibly proud of demonstrating true Sovereign Deployability. By proving that local 8B models can handle the bulk of standard user queries entirely on-device, we showed how enterprises can adopt the Apertus models securely, without leaking sensitive PII to external APIs, while drastically cutting inference costs.
What we learned
We gained deep insights into the Apertus model family architecture. We learned how to effectively quantize and serve the 8B model locally, and we learned how to build highly concurrent inference routers in Node.js that handle fallback logic gracefully.
What's next for EdgeRouter: Sovereign Apertus Gateway
We plan to open-source the router as a standalone NPM package so any enterprise in Switzerland (or globally) can drop it into their stack. Next, we want to add a Semantic Cache layer—so if the Apertus 70B model answers a complex question once, EdgeRouter will cache the embedding and instantly return it for future users without hitting the model again.
Built With
- artificial-intelligence
- css3
- dashboard
- data-science
- edge-computing
- html5
- javascript
- llama-cpp
- machine-learning
- node.js

Log in or sign up for Devpost to join the conversation.