Inspiration

CALL-E's own documentation is honest about a gap: completion_confidence is confidence that the task reached a clear end state — explicitly not confidence in the quality of the business answer. That sentence stuck with us. Every phone agent extracts what was said and throws away how the speaker knew it. We watched two hypothetical calls both answer "Tuesday" — one after "hold on, let me check" and twelve seconds of silence, one instantly with "should be Tuesday" — and realized the transcript already contains the difference. Speech carries the speaker's epistemic state; structured extraction discards it. We wanted to put it back.

What it does

provenance-grade attaches a knowledge grade to every field a CALL-E agent extracts, computed purely from the transcript_turns the API already returns: verified (the speaker demonstrably consulted something), asserted (stated directly, but nothing established where it came from), assumed (hedged, deferred, misaligned, or a bare round number), or unstated (the transcript never answered it — never a grade). Each grade ships with the exact supporting transcript span and the signals that earned it, so every grade is auditable. It also includes a pre-call lint that rewrites the task text so the signals are elicitable in the first place — forcing read-backs of critical values, turning "do you know?" into "can you check?", and asking for one corroborating specific. Downstream, the consumer rule is simple: act on verified, human-in-the-loop on asserted, never auto-act on assumed or unstated.

How we built it

Nine speech signals — retrieval gap, explicit check language, corroborating specifics nobody asked for, hedging, deferral, read-back compliance, round-number shape, answer alignment, and self-correction — feed a fixed, fail-closed rule table. Seven signals are pure regex and arithmetic; the two semantic ones (corroboration, alignment) run on a deterministic entity heuristic and expose provider interfaces where a model may return spans but never a grade. The grade is always computed by the rule table, which is what makes it testable. We wrote the 24 hand-labelled fixtures before the grader — voicemail, IVR, deferrals, self-corrections, read-back refusals, code-switched Hindi-English — so the eval harness existed before the thing it evaluates. TypeScript, zod, vitest; zero network in every code path; zero dedicated phone calls spent.

Challenges we ran into

The timing signal is structurally weak: CALL-E's offset_seconds is an integer marking turn start, so the "lookup pause" absorbs the previous turn's speaking time. We had to demote latency to the weakest signal and lean on check language and corroboration instead. Hedging detection is English-centric, and code-switched Indian business calls broke it immediately — "Tuesday tak ho jayega, shayad" is a hedged answer our English lexicon reads as confident. And the hardest problem was philosophical: keeping the grader honest. It is very tempting to let a model just "judge" confidence; we refused, because a model's opinion isn't testable and a rule table is.

Accomplishments that we're proud of

23/24 on the confusion matrix — and we're proudest of the 24th. The single miss is shipped deliberately as a fixture: the Hindi hedge our English lexicon cannot see, documented in the limitations file, sitting as the one red cell in our own demo video. A named limitation beats a silent one. Beyond that: 125 offline tests, a fail-closed default where absence of signal never upgrades a field, and an ethics boundary that is structural rather than aspirational — grades attach to the organisation, never the person, and the test suite asserts the output schema contains no field that could identify who answered.

What we learned

That restraint is a feature. The strongest skills in this ecosystem brag about what they refuse to do — fail closed, no network, block by default — and building in that culture changed our design: our default grade is "assumed" and everything above it must be earned. We also learned that you can improve a signal before you measure it: the pre-call lint exists because a bad grade is often the caller's fault — if you never asked for a read-back, you can't complain the value was never confirmed. And we learned that speech really does leak epistemic state in machine-readable ways: people who just read a screen volunteer what they saw.

What's next for Provenance Grade

Validation against reality. Every grade is currently a prediction with a rationale, not a calibrated probability: the next step is to record the grade, wait for the promised event, record whether it held, and measure grade-versus-outcome accuracy per organisation. Over enough calls that produces a calibration curve and a per-supplier reliability score — a supplier whose "asserted" Tuesdays actually arrive earns trust; one whose "verified" claims fail flags a broken lookup process. Nearer term: validated Hindi and Tamil lexicons with their own labelled fixtures (killing our one red cell honestly), and model-assisted span providers to widen recall on the two semantic signals — spans in, never grades out.

Built With

Share this project:

Updates

Submission history