-
-
An agentic verification engine: it calls the real world with CALL-E and returns evidence-backed claims — never a guess, and an honest no.
-
The goal compiled into hard requirements — each shown with the exact question CALL-E will ask on the phone. Checkable gates, not vibes.
-
Calls stream in: dialing, live transcript, structured result, then a verdict per constraint. Every strategist decision is logged live.
-
Click any verified claim: CALL-E call ID, confidence, and the supplier's own words — pending manager approval, then confirmed until 5 PM.
-
REALITY VERIFIED — 6/6 hard requirements, 93% confidence. Two suppliers were never called: the run stopped early once one fully verified.
-
When the internet isn't enough, ask the people who know. CALL-E does the phone work; GroundTruth does the constraint solving.
Inspiration
A technician drives forty minutes to a supplier whose website said "in stock." It isn't. The listing was stale, the compatible revision was the XZ-420B, and nobody could have known without asking a person.
That fact was never written down. Stock changes hourly, compatibility lives in a counter clerk's head, and a hold until 5 PM is a favour between businesses. No API, no scrape, and no model trained on last year's web can settle it. A phone call can.
So the interesting question isn't "can an AI make a phone call?" It's:
What does an agent owe you when the only source of truth is a stranger who might hedge, misremember, or say "I think so"?
What it does
GroundTruth compiles a plain-language operational goal into hard constraints, calls candidates through CALL-E, adapts its questions to what each person actually says, and returns claims with evidence — each linked to the CALL-E call id, the supplier's own words, and a confidence score.
It is not an AI phone caller with a memory. It's a constraint solver that uses CALL-E as a sensor.
Three rules do most of the work:
- A hedge is not a fact. "Probably", "I think", "we usually have it" are
uncertain. One polite follow-up is allowed; after that the constraint stays UNKNOWN and the system says so. - No evidence, no verified claim. Verification needs the validated structured result and stored provenance. Verified claims are immutable; contradictions are recorded, never erased.
- An honest "no" is a result. When nothing satisfies every hard requirement, the output is
NO FULLY VERIFIED MATCHwith a per-candidate breakdown — never a softened yes.
The part I'm proudest of: refusing to dial
A published CALL-E Goal owns a version-pinned resultSchema. It can be perfectly healthy and still be structurally incapable of answering one of your hard constraints, because no declared field can carry the answer. Run it anyway and you spend a real phone call on a real person to receive a constraint that can only ever be UNKNOWN.
So GroundTruth type-checks the Goal against the constraint set before the run is created. Every phone-derived constraint must bind to a declared result field; every required input variable must be suppliable. If not, the task fails with GOAL_INCOMPATIBLE and no call is placed.
It's the project's core rule — never claim a verification the evidence cannot support — moved one step earlier, to before the call exists.
It ships as a zero-dependency script anyone can use:
node scripts/check-goal-compatibility.mjs goal.json --spec published-goal.json
# exit 0 = compatible · exit 2 = do not run this Goal
INCOMPATIBLE — "Example stock check" v2 cannot verify this goal
(no declared result field for: compatibility, pickup_today, hold_until).
Do not run it: the constraint would stay UNKNOWN after a real call.
Proof: real calls, and an honest no
I placed three real verification calls through CALL-E to a consenting test number, playing the supplier myself. Two of them ran clean, and each returned four verified claims with evidence — ₹22,800, two units in stock, same-day pickup, compatibility confirmed — at 0.86 and 0.92 completion confidence. Both times the hold was correctly left UNKNOWN, pending a manager.
The first is the one worth showing. call_DxfIoUheBQMhn2GeKu0_Ww, 130 seconds, 37 transcript turns, completion confidence 0.78.
The supplier was evasive the way real people are — a talk-over, a "regarding what are you discussing about?", half-answers. The first stock answer was hedged, so the agent fired its planned fallback question and got a definite answer:
[107s] GT : Thanks — is a genuine XZ-420 compressor
physically in stock right now, or not?
[115s] SUP: Correct. No, not at the moment.
[118s] GT : Could you tell me whether pickup is available
today, and if so, what time?
[126s] SUP: No.
Five phone-derived constraints. Two settled by a definite answer. Three left open:
CONSTRAINT RESULT CLAIM
in stock not_available failed
pickup today false failed
price never answered unknown <- not guessed
compatibility uncertain unknown <- not promoted
hold until 5 PM never answered unknown
DECISION: no_match, confidence 0
Not one fact was invented to manufacture a result. That is the whole thesis, on a real phone call with a real human — and it's a better demonstration than a success would have been.
How I built it
Next.js 16 and TypeScript, with CALL-E behind a single adapter seam that has three implementations: ad-hoc call tasks (calls.create), published Goals (goals.run), and a deterministic mock for the offline demo. The orchestrator never knows which one it's talking to.
- Verification core — deterministic constraint evaluation with no LLM in the loop; a claim lifecycle of
unknown → verified / contradicted / failed / expired; confidence that never overrides verification state. - Adaptive strategist — decides follow-up / next candidate / stop from persisted state. The second call to a supplier exists because the first left a constraint UNKNOWN, not because a script said so.
- Safety by construction — authorization gates at the orchestration layer, a negation-aware prohibited-phrase scanner on the exact text handed to CALL-E, PII redaction, masked numbers, and a full audit log.
- Storage — Drizzle + PostgreSQL, or an in-memory store, behind one interface.
- 79 unit and integration tests, 2 Playwright end-to-end flows against a production build.
The demo runs entirely offline against a deterministic mock, so it behaves identically every time and needs no credentials. That determinism is a feature, not a shortcut — it's what makes the thing testable.
Challenges
The headline feature that never worked. The adaptive follow-up — a second call generated because the first left a constraint UNKNOWN — is the behaviour I demo. On a real call the supplier said the exact sentence it exists for: "I can hold it, but I need to check with my manager first." No follow-up happened. The trigger required hold_available: true; CALL-E's real extraction returned hold_available: null, hold_confirmed: false. Only my own mock ever produced that positive flag — so my tests were validating the code against the fixture that generated it, and the feature would never once have fired on a real phone call. It now keys on refusal: an explicit no closes the door, anything short of it is unresolved. The same call taught me a second thing — a follow-up dialled ~90 seconds after hanging up failed at the provider, and redialling someone that fast after asking them to go check with a manager was never realistic anyway. Follow-ups now wait.
The bug that contradicted the whole thesis. Clicking the "Hold confirmed until 5 PM" chip — the exact step my own demo script told judges to take — showed evidence from the first call, reading "We can hold it, but I need to confirm with my manager first." A verified claim, backed by words saying it wasn't confirmed. Two stacked defects: a membership guard that skipped the store write, and evidence records addressed to a draft's id so the UI couldn't group them. Now covered by a regression test I verified fails against the old code.
A crash only production could show me. Running a production build against a real database, every request 500'd — TypeError: i is not a constructor. The bundler had rewritten pg's exports. It reproduced only with DATABASE_URL set and only in built output: invisible to next dev and to all 75 tests passing at the time, and exactly the deployed configuration.
A safety hole I created by accident. Every task seeds itself with demo supplier personas whose numbers are fictional but well-formed Indian landlines. Nothing stopped that in real mode — a live run would have dialled six strangers. The provider now refuses outside mock mode, and the real-call harness will not start unless the pending candidate set is exactly the one number you nominated.
My own secret scanner didn't match real keys. It was written against the calle_live_* shape in the SDK's examples. Production keys are iams_live_…. A real key would have been stored and displayed verbatim.
What I learned
Every one of those defects survived a green test suite. They were found by running the thing — in production configuration, against a real database, on a real phone call. A verification project that hadn't verified itself would have been the wrong kind of ironic.
The follow-up bug is the one I keep thinking about, because the failure was circular: the mock defined a field's shape, the code was written to that shape, and the tests confirmed the code matched the mock that produced it. Nothing in that loop touches reality. Determinism makes a demo trustworthy and a test suite fast, and it will also happily certify a feature that cannot work.
I also learned how much of the honesty has to be structural. "Don't fabricate" as a prompt instruction is a wish. As a claim lifecycle that refuses the unknown → verified transition without attached evidence, it's a guarantee. The same goes for safety: a prompt saying "don't buy anything" is weaker than an authorization set that never contains a purchase action, plus a scanner on the exact text handed to CALL-E.
And the small things stuck with me. The speech recogniser heard "very far" as "very fast" — on the distance question, precisely where it mattered. The voice read the part number aloud as "capitalized X, capitalized Z, four, two, zero", which in parts sourcing is the single most important string in the call. Later attempts failed before the phone even rang, with opaque provider codes and zero duration. Reality is noisy in ways a mock never is, and that is the argument for treating every hedge as UNKNOWN rather than parsing it optimistically.
What's next
- Publish the verification protocol as a CALL-E Goal so the contract itself is shareable and version-pinned.
- A real discovery provider — candidate lists are demo/manual today by design.
- Expiry as a first-class citizen: stock and holds go stale, and a claim verified an hour ago may not be true now.
The verification protocol is packaged as a reusable Agent Skill and submitted to awesome-phone-call-agents, with two dependency-free scripts so anyone can compose a protocol-compliant call task or check a Goal's contract before spending a call on it.
Built With
- call-e
- drizzle-orm
- framer-motion
- neon
- next.js
- playwright
- postgresql
- react
- tailwindcss
- typescript
- vercel
- vitest
- zod
Log in or sign up for Devpost to join the conversation.