Inspiration

Every agent hackathon has the same failure mode: agents make claims, and other agents have no way to check them. In the Q&A room before the event, we watched agents argue about versions, APIs, and pricing . Some confidently wrong, some confidently right, nobody able to settle it. The services in the room that offered "verification" all had the same weakness: they'd answer even when they had nothing to check against.

We wanted to build the opposite. A service whose default response to uncertainty is silence, not a guess.

What it does

Arbiter has two endpoints and one guarantee.

POST /service/verify takes a claim and returns one of three verdicts — supported, contradicted, or insufficient — plus the live web sources that justify it.

POST /service/research takes a question and returns a grounded answer with source URLs, or an explicit refusal.

The guarantee is the same for both: if search returns nothing usable, the service refuses. It does not answer from model memory. The model never sees the claim unless search returned results, so it cannot invent one.

Arbiter is also autonomous in the Arena room. Any agent that types "verify: " gets a grounded verdict reply within five seconds, posted back to the room. No human relay.

How we built it

A Node.js + Express server exposes the two endpoints. Each request runs through the @aicoo/sharedos kernel before anything else happens: a grant must authorize the turn, one purpose string must match, deny-by-default means nothing else is reachable. Only after the kernel admits the turn does the service run.

Inside the turn: DuckDuckGo Lite (keyless, GET-based) is the primary search source, with the Wikipedia REST API as a fallback. Retrieved snippets are formatted into a strict prompt that instructs the model to answer only from what it was given, or reply with a specific refusal string. Groq's openai/gpt-oss-20b does the synthesis.

The "sources" field in the response is built by extracting citation markers like [1] and 【2】 from the model's output — so only sources the model actually cited appear, not every article the search returned.

A separate poll.js script watches the SharedNet room for messages matching verify: , calls the local service, and posts the verdict back. This makes the whole loop autonomous.

Challenges we ran into

The biggest one was honesty under pressure. Our first version answered every question from model memory. When we asked it about the SharedOS npm package, it confidently replied "version 2.0.0" — a number that does not exist. That was the moment we committed to the refusal design: if search returns nothing, the model never runs.

The second was the environment. The SharedNet CLI refused to run on Windows because it checks for Unix-style owner-only credential permissions, which MINGW64 cannot provide. We ended up bypassing the CLI entirely and driving SharedNet through its HTTP API with curl and Node — which the npm docs confirm is a fully supported path.

The third was search reliability. DuckDuckGo's HTML endpoint throttled us and threw ECONNRESET. Switching to the lite.duckduckgo.com GET endpoint with a real browser User-Agent solved it; Wikipedia REST as a fallback covered the rest.

Accomplishments that we're proud of

The refusal behavior actually works, and it's a feature, not a limitation. Our service will say "No verified sources found" in a room full of agents that will not. We tested it live in the Arena Q&A room and it answered a claim from another seat autonomously, with a citeable source, in under five seconds.

We also kept the sources honest. Early versions returned every article search had matched — including irrelevant ones like the Eiffel Tower in Texas when the claim was about the Paris one. We fixed that by extracting only the citation markers the model actually used, so the sources list is exactly what justified the verdict.

And we got the whole loop running end-to-end on a public URL: Cloudflare tunnel, SharedOS kernel, live search, Groq, autonomous room reply.

What we learned

The hard part of grounding isn't the search — it's the discipline to not answer when the search fails. It is much easier to write a prompt that says "answer if you can" than to write code that returns nothing when there's nothing to return. The refusal path is where the integrity of the product lives.

We also learned that "sources" is not a list of links the search returned. It's a list of links the model actually cited. Those are different, and the difference matters if you care whether the output is honest.

On the infrastructure side: cross-platform CLI tools assume Unix. On Windows, driving the same service over plain HTTP turned out to be faster, more reliable, and easier to debug than fighting the CLI's permission model.

What's next for Arbiter - Grounded Claim Verification

Three things.

First, tighten the search. Add a second keyless source so the fallback path is broader than Wikipedia, and cache query results per claim so repeated claims cost nothing.

Second, add a "confidence" field. Right now the verdict is binary-categorical — supported, contradicted, insufficient. A confidence score derived from how many independent sources agree would make the output more useful for arbitration between two agents that disagree.

Third, wire in the payment and receipt flow properly. Right now Arbiter is free in the room. A real product needs per-call settlement and a signed verdict receipt the buyer can attach to their own defense. That turns Arbiter from a service into infrastructure for agent disputes.

Share this project:

Updates

Submission history