Inspiration

We went looking for a business problem that was real rather than theatrical, and found one hiding in a single sentence of Australian law.

Under section 72 of the National Credit Code, a borrower gives a hardship notice orally. There is no form. They never have to say the word "hardship". They say "I can't make the repayments, not for a few months", and the lender then has 21 days to assess it and notify a decision. The clock starts when the customer speaks, not when anyone writes it down.

Which means the entire statutory obligation rests on whether the person on the phone recognised it.

Usually nobody does. In August 2025, ASIC penalised NAB and AFSH Nominees $15.5 million (25-165MR) because 345 customers gave notice and never received a response within the required period, over conduct running from 2018 to 2023. Nobody in those calls decided to break the law. They just didn't hear the clock start.

What it does

CallFlag listens to a hardship call as it happens and does three things:

  1. Names the duty that just triggered.
  2. Shows the date it falls due, computed from a rule table.
  3. Gives the worker the one question that settles it.

When the call ends it produces a report card: what was caught, what was missed, what is still owed, and which clocks are now running.

It never speaks to the customer and it never makes the decision. It puts a duty, a date and a question in front of a human.

How we built it

React, TypeScript and Vite on the front end, Supabase Edge Functions (Deno) on the back, Gemini for the language the rules can't reach, ElevenLabs for practice mode.

The decision the whole thing rests on is that the threshold scales with the consequence. Events come in three kinds:

kind Bar Asserts Persisted Enters the compliance record
obligation High, inability established A legal duty is owed Yes Yes
request Low, a hint or an ask The customer asked for a change Yes No
cue Engine's own A circumstance that routes the call Never No

A prompt to ask a clarifying question can fire on a hint, because a false positive costs one question. A claim that a statutory clock started has to be earned. Deadlines come from a rule table in code, never from a model.

Detection runs in two stages. A deterministic lexicon looks for both halves of the s 72 test, payment inability and a medium-term horizon, and a recovery signal anywhere nearby suppresses it: "I'll be square once the 20th lands" is a timing problem, not hardship. Only genuinely ambiguous turns reach the model, which is asked one question through a response schema that has no field for the cause. It cannot return "illness" or "job loss" because there is nowhere to put it.

Challenges we ran into

Attribution turned out to be load-bearing. Only a customer turn can raise an obligation and only a staff turn can discharge one, and in live dictation on a single microphone the speaker is guessed by a model. So the same rate limit that degrades the AI also corrupts the transcript. We added a confidence tier: a turn whose speaker was only inferred can prompt but can never assert. An inability stated on an inferred turn becomes a request, not a clock.

A checkbox is not proof. Overlapping or unattributed speech cannot support a compliance verdict. CallFlag links every judgement to timestamped staff evidence, and where the transcript cannot prove the response the result stays unverified and no score is awarded. We also separated unresolved hardship requests into an unscored seven-day follow-up, so an operational callback is never presented as a statutory deadline.

Our own evaluation lied to us twice. Once because a cross-check script ignored which signs were already on record and scored correct detections as false positives. Once because our duration patterns were overfitted to the four calls they were written beside, and matched zero of an independent set. Both were found by teammates reviewing each other's code.

And the honest limitation. Our rule requires both halves of the s 72 test in the same customer turn. Real callers split them across three: the inability early, the duration much later. Measured against our own demo script, the deterministic pass raises nothing at all. So the claim we make is the narrow true one: clear statutory notices can be raised locally without an API call, while the model handles less explicit language.

Accomplishments that we're proud of

The A/B replay, and specifically the way it refuses to lie. The same call runs twice, once with coaching on screen and once silent. Both runs produce an identical event stream, because the detector cannot see the switch. The build script that generates the comparison page will not write it unless the two transcripts are byte-identical up to the moment they diverge and the silent side displays nothing. The difference on screen is the worker's chance to act, not a different analysis, and that is enforced by a test rather than promised in a pitch.

Also: 43 of 44 labelled cases agree with an independently written engine, with zero false positives, in rules mode, with no API calls at all. And four real bugs caught in peer review and fixed with measurements rather than assurances, including one that would have let a single network failure take down detections the deterministic pass had already found.

What we learned

That the hard part of applied AI compliance is not detection. It is deciding what you are willing to assert. Every interesting argument we had this weekend was about where a hint becomes a legal claim, because getting it wrong in one direction misses a notice and in the other invents a duty nobody owes.

We also learned to write down what we could not prove. The repository says which cases we still disagree on, which claims hold only in one mode, and that no connected ElevenLabs voice session was recorded end to end because a free-tier minute quota ran out. A number you can't stand behind is worth less than a limitation you can.

What's next for CallFlag

Cross-turn accumulation of the statutory test, so the two halves can arrive in different sentences the way people actually speak. Dual-channel telephony audio, which is the only real fix for attribution. Self-hosting inside a lender's own boundary.

And the deployment shape we think actually sells: run it silent for a month. It listens, judges and starts the clocks while saying nothing to anyone. The bank gets a report card per call and a deadline list, and finds out what it has been missing before a single AI word reaches a customer. Then they turn coaching on.

Built With

Share this project:

Updates

Submission history