Inspiration

Getting four quotes for the same job is a chore everybody knows and nobody does. You find the businesses, you check they are open, you call them one at a time, you write the prices on the back of an envelope, and by the fourth call you have forgotten what the first one said.

We went looking at what already exists in the CALL-E apps repository and found that all ten apps share one starting point: a list somebody already had. A fixture, a CSV, a study file, a hand-written plan. hungrycall-cascade even says so out loud — its restaurant search lives in a different repository, outside the contribution.

So the interesting question was not how do I call these ten businesses? It was the one nobody had answered: who should I even call?

What it does

QuoteRunner turns an intention — "replace the windscreen on a 2019 Honda Civic" — into a price table.

It screens candidates down to the ones that publish a number, are open right now, and can actually be reached, deriving each calling window from published opening_hours rather than from configuration. Then it calls them through CALL-E with a typed result_schema, and sorts what they said by price.

       price  available    warranty   business
------------  ------------ ---------- ----------------------------
  199.50 USD  2026-08-14   6 mo       Riverside Motors
     245 USD  2026-08-11   12 mo      Northgate Auto Glass
 280-320 USD  2026-08-10   24 mo      Halloran & Sons Bodyshop

Cheapest: Riverside Motors

Answered, no price given:
   - Quickfit Windscreens  ->  Service manager was out; asked to be called back tomorrow.

Exclusions are output, not a silent filter. A run that quietly dropped half its candidates looks identical to one that found nothing worth calling, and the difference matters to whoever reads the result.

The example is a windscreen quote. The pattern is not: three plumbers with a free slot this week, five suppliers who stock the part, the clinics within range open on a Saturday. Same mechanism, different subject.

How we built it

Two layers, deliberately separated.

Screening has no dependencies and never imports the SDK. It parses OpenStreetMap opening_hours, validates numbers as E.164, and decides who is callable. Provider-agnostic, and testable without credentials.

Execution places the calls through the calle-ai SDK with a nine-field result_schema. Every field is a string with an explicit unknown, including the price — a receptionist who says "depends on the glass, call back Tuesday" is a normal outcome, and a numeric price field would force the model to invent a number to satisfy the type.

Three modes, and only one dials: preview is the default and reads no credentials, --simulate runs the whole pipeline against canned answers, and --execute is gated four ways.

The part we are most pleased with

The confirmation token.

--execute needs --confirm <token>, and the token is a hash of the job plus the sorted list of numbers. A token you obtained by reviewing one list will not authorise a different list. Re-plan an hour later, get different candidates because a shop closed, and the old token stops working.

That is not decoration. It is the gap we reported through this hackathon's own feedback form: call start has no machine-enforced confirmation, so the only thing standing between an agent and a live call is a sentence in a markdown file that a model is free to skip. A hash the operator has to paste back cannot be skipped by a model that is feeling confident. We found the hole, reported it, and then built the fix into our own app.

The fourth gate is quieter and matters just as much: opening hours are re-checked at dial time, not at plan time. A batch of twelve calls takes minutes, and a shop that closes at 18:00 must not be dialled at 18:04 because it was open when the plan was written.

Challenges we ran into

A date is eight digits with separators. We redact phone numbers and emails out of everything before it is stored, because evidence_summary is model-written prose repeating what a person said out loud and can contain anything they read out. The first version applied that redaction to every field — and quietly turned every availability date into [number redacted]. The fix was to stop using a blunt instrument: narrow fields are validated against the shape they are supposed to have and become unknown otherwise, which also catches a phone number smuggled into a date field.

Comparing prices you cannot compare. If quotes come back in more than one currency, QuoteRunner reports them and refuses to rank them. Converting them would invent an exchange rate nobody quoted. And an unparseable price sorts last as "no price given" rather than becoming a zero that wins the comparison.

Deciding what not to build. Our first version led with a safety layer, until we read consent-gate properly and found it far more rigorous than ours. Leading with safety was asking to be compared against something better made. So we cut it back to a requirement met rather than a headline, and QuoteRunner now composes with ConsentGate instead of competing: ConsentGate requires the caller to supply "a recipient timezone and permitted calling window"; QuoteRunner derives that window from open data.

What we learned

opening_hours is an ordinary OpenStreetMap tag. Treated as decoration it tells you when a shop is open. Treated as a control it decides whether a call may be placed at all — and it is already published, in a machine-readable format, for millions of businesses that nobody has to onboard.

That is what makes the discovery layer worth having. Not that it finds businesses, but that the same open data that finds them also says when it is acceptable to ring them.

What's next

Wiring the OpenStreetMap discovery front end back in, so the input is a niche and a place rather than a fixture. It exists and runs — it produced about 1,980 real businesses across three runs while we were calibrating — but the fixture is what keeps the contribution runnable with no network and no credentials, which is what the repository asks for.

After that: callbacks. Roughly one business in four asks to be called back at a better time, and right now that is recorded and left alone rather than rescheduled — on purpose, because a redial the operator did not ask for is a second call to a real business.

Tests

Ran 127 tests in 0.239s
OK

No test places a call or reads a credential. The CallsAPI is a fake, which is the point: the gates that stop a real call have to be testable without placing one.

Built With

Share this project:

Updates

Submission history