Who this is for
Somebody about to commit to an offer whose page will not settle the one condition that decides it. A clause that says entrants ranked outside the top ten pay a fee to reach the second stage. A tier described as free whose footnote mentions a charge. A programme open to everyone except for four words beside the register button.
They are not lawyers and not procurement. They are one person deciding whether this particular thing is worth their week, holding a page that answers every question except the one that matters. The page is not lying. It is silent, and silence reads the same as a yes until it costs you.
Today that person has two options and neither is good. Commit and find out, or call and ask, which means finding the number, waiting, and spending a stranger's minute on a question that a page could have answered. Most people commit and find out, which is how the clause pays for itself.
The reason this is worth running after the hackathon is that the shape recurs. Any offer with a condition that only a human can confirm produces the same moment, and the tool does not need to know the domain to place the call. What it needs to know, and what the rule at the heart of it enforces, is when the call is not worth placing at all. That is the part that makes it usable more than once without becoming a nuisance to the people on the other end.
Inspiration
I build a tool that reads an offer page and reports where it contradicts itself. Free that becomes charged three sections down. Cash prizes that a footnote says are not cash. Open to everyone, except for four words beside the register button.
It works, and it stops at the edge of the real world. A page is a document. Documents can be wrong in a second way that no amount of reading catches.
A page can contradict itself. It can also contradict the person who answers the phone.
The moment that made this concrete. I had read the rules of a contest and found a clause saying that entrants ranked outside the top ten would pay a fee to reach the second stage. That single sentence changes whether the thing is worth entering. Nothing on the page confirmed or denied it beyond that sentence. The only way to settle it was to ask a human being, out loud, and the cost of asking was a phone call I was not going to make by hand for every page I read.
What it does
It takes one finding from the page reader, the condition an offer sets quietly, and it turns that finding into a call worth placing.
a quoted clause from the page
-> a call task in plain language
-> a result schema the provider fills from what was said
-> the page's words and the caller's words, side by side
The output is never a verdict. It is two quotations, the exact sentence from the page and the structured answer from the call, next to each other. The reader decides.
The rule that governs the whole thing
Never make a stranger repeat on the phone what the page already says.
A call is worth placing only if its answer can CONTRADICT the page. A question whose two possible answers both leave the matter where it stood is a question you do not ask, because asking it spends a minute of someone who did not volunteer.
That rule is not a comment. It is enforced. Six clause families are considered
callable. Anything outside them raises NothingToAsk, which is not an error
and is the most common outcome on real pages. Run the dry-run path against a
bank's pricing page and it prints, correctly, that no call is justified.
Three constraints the witnesses enforce
The call declares itself, always. A synthetic voice that does not say what it is is a deception, and that would be a strange thing to build into a tool whose entire subject is what people hide from you.
One question per call. The task says so twice, because a model given room will fill it.
unknown is always available, and it is never a contradiction. Without a
way to say the call settled nothing, the extraction model has to pick between
yes and no when the call produced neither, and it will pick. An absence of
answer is not a denial.
The part I would want a judge to look at
Two things in this contribution came from reading the provider's own contract rather than from building faster.
A malformed result schema is not refused when the call is created. The
call is queued, dialled, answered, and only after it reaches a terminal state
does extraction run against a schema that was never going to work. You have
spent a call, a share of a twenty-call budget, and a stranger's minute, to
learn nothing. So validate_result_schema compares the schema against the
feature list the provider publishes, before anything is dialled. One of its
rules is stricter than the contract, and it is deliberate. An enum with no way
to say the call settled nothing is refused outright.
A question nobody can answer is not asked. One of my six families used to send an open question, which countries of residence are accepted, while expecting a yes or no field back, without ever naming the country at stake. The extraction model would have received prose and a binary field with no reference point. It would have returned a value, and that value would have been invented. Families that need a piece of context now refuse to produce a call task without it.
Neither of those is a feature. Both are ways of not producing a confident answer out of nothing, which is the failure mode that matters when the output is going to be read as evidence.
How it was built, and what it cost
It calls CALL-E at runtime. place_call.py posts to POST /v1/calls with
a Bearer key read from the environment, and reads the finished call back from
GET /v1/calls/{id} to turn the structured result into a contradiction, or
into nothing. One real call was placed against that path, to a number I had
explicit authorisation to dial. It connected, ran eighteen seconds, and proved
the chain end to end.
Everything that can be settled without dialling is settled without
dialling, and that is a separation, not an absence. The module that builds
the request opens no socket and knows no key, so the part that costs a real
call from a budget of twenty is the only part you cannot exercise for free.
Four refusals sit in front of it, and none of them costs a call. A clause
outside the six families never becomes a request. A recipient the operator has
not authorised is refused even when the number is well formed, because a valid
number is not the same thing as a number you are allowed to call. A schema the
provider would accept and then fail to fill is refused before dialling. And a
missing key stops the run with a sentence rather than with a 401 that reads
like a permissions problem.
The whole thing is held by 56 offline witnesses, twenty-seven on the file that prepares and twenty-nine on the file that dials. Six of them guard the schema validator, five built to be refused and one to check that every family still emits a conforming schema, because a validator that only ever refuses is as useless as one that only ever accepts. The transport is a parameter, which is what lets the witnesses exercise the dialling file without a key and without ringing anyone.
A reviewer caught me too, and that one is worth more than the tests. A
maintainer read the pull request and found three things. The recipient pattern
took a country code starting with zero. The runnable snippets used a
placeholder the bridge itself rejects, while the README claimed every sample
used a fictional number. And the README said the contribution contained no
client, when the file that dials is a credential-bearing HTTP client that can
ring a real person. The third one was a materially false claim about my own
work, in my own words, and no test would ever have found it. All three are
fixed, the recipient now has to be listed in CALLE_ALLOWED_RECIPIENTS before
anything is dialled, and the samples use a number inside the block Ofcom
reserves for drama.
And one of those witnesses caught me. The nine last tests in the first
file never ran, because the __main__ block had ended up in the middle of it
and unittest.main() exits. Fourteen of twenty-three were running, and the
nine missing ones were the hostile schemas, the half worth having. It is fixed,
it is a separate commit, and I mention it because a bench you have not counted
is a bench you do not have.
What I did not do, and why it belongs here
I did not call a single company that had not asked to be called. The honest demonstration of this tool would be to ring a real vendor and ask them about a clause on their own page. It would be more impressive and it would spend the time of people who never volunteered. The demonstration uses a number I am authorised to call, which makes the demo weaker, and the video says so out loud.
There is no sandbox on this platform, no echo number, no simulated call. I learned that by needing one. It is the first item in the feedback document attached to this submission, along with an API key list that shows a masked value long enough to look complete, and a dashboard that hides what the machine contract answers in one file.
Built with
Python, no dependencies, standard library only. The provider's REST API,
POST /v1/calls and GET /v1/calls/{id}, and its published OpenAPI contract,
which answered more questions than the dashboard did. 56 unit tests, no
network, no key. A no-call path, which the target repository asks every
runnable contribution to provide, and which turned out to be the thing that
made the project buildable at all.
Links
The contribution was merged into the provider's public repository,
CALLE-AI/awesome-phone-call-agents, as pull request #264 on 2026-09-05, and
lives under apps/python/clause-check-by-phone. It passes the repository's own
scripts/validate_repository.py on a fresh clone.
A post-merge review found two more real defects and both are fixed in pull request #330, merged on 2026-09-06. The bounding helper that keeps the destination and the transcript out of what a caller prints only removed top-level fields, so a payload shaped the way the repository's own fake server shapes one came back whole, and the create request omitted the recipient, which made the live path unusable. Seven more witnesses came with the fix.
Two smaller pull requests came out of running the repository's suites on a
machine that is not the author's, #331 and #332, both merged the same day.
Neither touches an app, and neither is mine. One pins a fixture's timezone, the
other guards two 0600 file-mode assertions behind POSIX. Three suites failed
on main outside North America or on Windows before them, and not one of those
failures said anything about the code.
Every number a reader is invited to run is +447700900123, inside the block
Ofcom reserves for drama, 07700 900xxx. It is a valid E.164 number, so the
snippets run exactly as written, and it is reserved, so it reaches nobody. One
more reserved number, +447700900999, appears in the witness for a recipient
the operator has not authorised. Three further values live only inside
witnesses for the number format itself, 00 00 00 00 00 and +0123456789,
which are not numbers at all, and +1234567, which tests the shortest length
E.164 allows. None of the five reaches a person.
That paragraph used to say something else, and the change is the reviewer's doing. The samples were all zeros, which reads safe and is not a valid number, so the runnable snippet in the README was rejected by the code it was demonstrating. A safe-looking example that does not run is worse than no example.
Built With
- call-e
- openapi
- python
- rest-api
Log in or sign up for Devpost to join the conversation.