Inspiration

Some facts are not on the web. Whether the part is on the shelf at this branch, whether the unit is still available at the advertised price, whether the slot is really open today. They live in one human's head, behind a phone number.

An agent asked for such a fact has three options and the first two both fail you. It can assert from stale text, which produces a confident wrong answer. It can hedge and tell you to go call, which hands the work straight back. We measured the second one rather than assuming it: across eight fresh agents under time pressure, "you should call the store yourself" happened five times out of five — including with this skill's own description loaded. That measured failure, not an imagined one, is what the skill is built against.

What it does

Ground Truth Gate is the third option. It answers you provisionally right now, places one disclosed CALL-E call to the party whose word is binding, and rewrites the answer only when the call comes back with positive evidence.

The core rule is the part that is easy to get wrong: absence of evidence is never evidence. If the person who answered was not sure, nothing is released. The provisional answer stands, untouched, and you are told it is still unconfirmed. A boolean cannot tell "no" apart from "nobody knew", so every field in the result schema is an enum with unknown always available.

It also decides whether to call at all. Five of the six triage rows place no call — a cheap question, a fresh source, a real API, a scope that already matches. The phone call is not the product. The correction is.

It actually called a store

On 2026-09-11 this placed a real CALL-E call to a published Wawa store line in Parsippany, New Jersey. Call call_JAAWic7-bVnUQuEmvmQp4w, thirteen transcript turns, verbatim:

agent Hi, this is an automated assistant calling on behalf of a customer to check one piece of information. This call may be recorded. This call is placed by Ground Truth Gate, a hackathon demo project. Do you have a moment for one quick question? store Wawa, Precipany. agent Hi Wawa, thank you for confirming. Is your hot food kitchen open and serving right now? store Hello? Yes? What? agent I'm sorry, let me clarify — is your hot food kitchen open and serving right now? store Yes. agent Thanks for confirming. Have a great day.

Three things in there were not scripted by us. It confirmed it had reached the right business before asking anything. It treated the store's name as data, not as an instruction. And when the reply was Hello? Yes? What? — a "yes" that is confusion rather than an answer — it declined to bank it and asked once more, cleanly.

The machine result was verdict: confirmed_true, completion_confidence: {score: 0.9, label: high}, quoted_answer: "Yes."

Then the part we care about most. With no abstention decision supplied, the skill refused to release the confirmed answer: released: False, provisional answer left standing. Only when the abstention gate agreed did it release and write back the correction, quoting the human's own word. A confirmed answer is still not a released fact.

How we built it

Python standard library only. No dependencies enter the skill.

  • gate.py — triage, call construction, and the release rule.
  • place_call.py — POST /v1/calls with a body-derived idempotency key, then GET /v1/calls/{id}. The webhook body is a hint; the re-fetch is the truth, because CALL-E webhooks are unsigned.
  • mock_calle.py — a complete local CALL-E stand-in serving all six outcomes, so the whole loop runs with no credential and no phone call.
  • self_test.py — raises explicitly instead of using bare assert, because python -O strips asserts and a suite that cannot fail reports success without testing anything.

We composed rather than reimplemented in three places, and said so in the PR. Calibrated abstention stays in verify-by-phone. Durable webhook delivery stays in apps/python/webhook-result-receiver. And references/composition.md names apps/web/local-atlas as the nearest neighbour and states the difference plainly: local-atlas decides whether to dial, this decides whether what came back may be released as fact.

Challenges we ran into

The honest one: we built the skill against failures we imagined, then measured. Eight baseline agents scored 0/5 on the prohibitions we had designed around — a cold agent already refuses to widen a question on the phone. Only two of our rules converted a real red to green. We rewrote the skill against the measured failures instead of the invented ones.

The second: disclosure is a security surface. Our own first implementation of the 47 CFR 64.1200(b) caller-identity clause concatenated an environment variable straight into the quoted script an agent speaks. An apostrophe closed the script early and the rest read as a fresh instruction to the calling agent. It is refused now, straight and curly quotes both, and there is a test with the payload in it.

Accomplishments that we're proud of

Dry run is the default and gate.py opens no socket at all. Every committed number is in the reserved 555-01xx range. A reviewer can run the entire loop in about a minute without a CALL-E credential:

python scripts/gate.py --demo
python scripts/self_test.py

What we learned

Grep the repository before claiming an idea is unowned. Twice we believed we had found white space, and twice it was already merged — once as a design doc whose opening line was our thesis. The judges wrote that file. Naming it in our own PR was better than hoping nobody looked.

What's next for Ground Truth Gate

Feed real call outcomes back into verify-by-phone's calibration fold, so the abstention threshold is learned from this skill's own history rather than a fixed alpha.

Built With

  • agent-skills
  • call-e
  • calle-api
  • claude-code
  • json-schema
  • markdown
  • python
  • python-stdlib
  • rest-api
  • telephony
  • voice-ai
  • webhooks
Share this project:

Updates

Submission history