As a solo developer, I already do this manually — copy the same hard question into ChatGPT, then Gemini, then Claude, then DeepSeek, and mentally compare the answers before trusting any single one. Every serious builder I know does some version of this, because no single model gets everything right, and blind trust in one AI's output is a real risk for technical, financial, or health-related decisions. But the process itself is tedious: opening tabs, retyping the same question, scrolling between windows. Consensus is the tool I wished existed — the workflow I was already doing, automated and made instant. What it does . . . Try it live: https://consensus-zeta-olive.vercel.app (password: cns-ZE8KszZih1hB). Running on a limited demo budget, so please don't overuse it — thank you! . . . . Consensus lets you ask one question and send it to multiple open-source AI models simultaneously — including NVIDIA's Nemotron via Nebius Token Factory. Every selected model streams its answer live, side by side, so you can watch them respond in real time instead of waiting on one model at a time. You can also pick a "Summarizer Boss" — one model that reads every other model's answer and produces a clear breakdown of where they agree, where they disagree, and what the most reasonable conclusion is. The whole point is turning "which AI do I trust here?" into "here's what several independent models actually think, and where the real uncertainty is." How we built it The backend is a lightweight Python FastAPI service that talks to Nebius Token Factory's OpenAI-compatible API — swapping models is just changing a model ID, not rewriting integration code. The frontend is a single HTML page with vanilla JavaScript, deliberately framework-free, so there's no build step between writing code and seeing it live. Each selected model is queried in its own parallel request and streams back token by token, so a slow model never blocks a fast one from showing its answer immediately. The whole thing is deployed on Vercel, with the API key and access password kept strictly server-side as environment variables — never exposed in the client or committed to the repository. Challenges we ran into The biggest challenge was cost safety on a genuinely small budget — this was built and tested on roughly $1 of API credit, as a solo, self-funded developer. That meant building real guardrails from day one: a password gate so the public link isn't discovered and drained by strangers, a hard per-response token cap so no single request can run away with the budget, and a live "max cost" estimate shown before you even hit send, so you always know what you're about to spend before spending it. Reasoning models added a second challenge — some return a long internal "thinking" block before the actual answer, which meant designing a collapsible thought-process section so the real answer isn't lost inside it. Accomplishments that we're proud of Getting true parallel streaming working cleanly — every model card updates independently and in real time, with no artificial waiting — was harder than it sounds, especially while keeping the codebase framework-free. I'm also proud that the safety mechanics (password protection, hard token limits, live cost estimation) aren't afterthoughts bolted on later; they were part of the architecture from the first version, because I built this the way I'd actually want to use it myself, on my own limited budget. What we learned Working directly with Nebius Token Factory's model catalog taught me how much variation exists even among open-source models in the same weight class — in response style, latency, and reasoning behavior. It also reinforced something I already believed as a developer who cross-checks AI outputs by habit: the real value isn't finding "the best model," it's seeing where independent models disagree, because that disagreement is often the most honest signal that a question doesn't have one clean answer. What's next for Consensus Next is a persistent history so past comparisons aren't lost on refresh, and letting users save a "model panel" they reuse often instead of reselecting every time. I'd also like to add lightweight scoring — letting users mark which model's answer actually turned out to be right after the fact, building a personal track record of which models to trust for which kinds of questions over time. . To try it live: https://consensus-zeta-olive.vercel.app (password: cns-ZE8KszZih1hB)

Built With

Share this project:

Updates

Submission history