Inspiration
I was using Claude for one of my personal projects when I noticed I kept repeating myself to get a task done. Claude would confidently give me something that looked right but wasn't, like a detail it made up or code that didn't actually work, and I'd spend a bunch of extra messages catching it and steering it back. That got me thinking: what if something caught those hallucinations the moment they happened? Less AI slop, and less time and tokens wasted going back and forth fixing answers I shouldn't have trusted in the first place.
What it does
Reigns is a desktop companion for Mac that keeps Claude on course. A little horse (Mamu) or unicorn (Mia) sits next to the Claude desktop app and checks every reply as it comes in.
It looks for the usual ways AI goes wrong: citations and DOIs that don't exist, wrong facts, code that calls functions that aren't real, summaries that add details your document never said, and Claude caving when you ask "are you sure?" even though it was right the first time.
When something doesn't check out, the pet gets alarmed and a bubble shows each problem along with the evidence, like the paper search that came back empty or the real function signature. It also reads the problems out loud, with a cowboy voice for Mamu and a unicorn voice for Mia, in English or Spanish. If an answer is clean, it just says good job.
Clicking Fix it pastes a fair recheck prompt back into Claude. The prompt shares what we found but still lets Claude stand its ground if it was actually right. The pet only calms down once the next answer really checks out. Each chat keeps its own panic score, and it learns over time from confirmed mistakes and your feedback.
How we built it
The Mac app is written in Swift. It reads the Claude app through the macOS Accessibility API, so there's no browser extension and nothing intercepting your traffic. It draws the pet and the bubble, pastes Fix-it prompts into Claude, and plays the voice.
Everything it reads goes to a local Python engine built with FastAPI over a WebSocket. The engine uses a fast Claude model to pull the checkable claims out of each reply, then sends each claim to the right checkers, all running at the same time. We check papers against Crossref, Semantic Scholar and OpenAlex, facts against Tavily web search and Wikipedia, and code against the real installed libraries (without ever running the AI's code). Other checkers re-ask the question to see if the answers stay consistent, catch Claude caving to pushback, and compare summaries against the pasted source. A claim only gets flagged red when there's real outside evidence behind it, and anything uncertain shows as amber. That feeds a panic score that drives the pet's mood.
For learning we used MongoDB Atlas. Confirmed mistakes are saved with Voyage AI embeddings, so vector search can pull up similar past cases. A bandit learns which recheck prompts actually work, and "I disagree" clicks tune the alarm thresholds. We never store conversation text in the database. The voice is written by Claude and spoken with ElevenLabs, and each chat's history and score stay in a local SQLite file on your Mac.
Challenges we ran into
The biggest one was precision. A hallucination checker that cries wolf is useless, so we spent a lot of time making sure it only raises a real alarm when there's actual evidence, especially in long chats where claims show up without context.
Speed was another. Checking every claim against outside sources takes time, so we run everything in parallel and let the pet react before the full explanation is ready.
We also hit a bug where saving each chat's panic score in the background made our tests freeze now and then, and fixing it meant rethinking how saving worked. Getting the voice right was harder than expected too: it had to cover every problem without rambling or cutting off mid-sentence, and switch language or character without resetting anything.
With four of us building different pieces at the same time, we had to agree on shared message formats and folder ownership early so nobody broke anyone else's work. And when we tested on real Claude answers, Claude often refused to make up citations, so building a fair test set meant checking every claim against real sources by hand.
Accomplishments that we're proud of
We built a full working product: it watches Claude live, checks claims against real evidence, explains why, and helps fix the answer. We tested it on 22 real Claude answers where we checked every claim by hand. It caught 2 of the 3 real mistakes, had one false alarm, and took about 7 seconds per reply. It's a small sample and we're upfront about that, but it's real data, not examples we made up.
We're also proud that it never runs AI-written code, keeps your conversations off the cloud database, and has a personality people actually enjoy. Plus the engine has around 420 automated tests behind it.
What we learned
Checking for hallucinations only works when it's grounded in outside evidence. Just asking a model "are you sure?" doesn't cut it, and models will often cave even when they were right. We also learned that newer models often refuse rather than make things up, but they still slip in subtler ways, like mislabeling a scientific model or using a library feature that was removed. And we learned that being honest about your numbers matters more than a demo that looks perfect.
What's next for Reigns
We want to add a second AI model as a cross-check so it's not just Claude checking Claude, and fix the two detector gaps our testing found. After that: a bigger, independently reviewed evaluation, support for other AI apps, the browser and Windows, and team features like a shared memory of known mistakes.



Log in or sign up for Devpost to join the conversation.