C2C (Cancellation to Compensation) is a durable agent for airline disruption compensation claims.

Everything outside the agent is synthetic — no real airline is contacted, no real claim is submitted, and every legal-looking artifact is stamped SYNTHETIC DEMO — NOT FOR SUBMISSION — NOT LEGAL ADVICE.

Repo: https://github.com/yablokolabs/c2c-agent

The problem A passenger whose flight was cancelled is probably owed money and has a job.

Deciding whether a claim is worth anything takes a competent person ~15 minutes. Everything after that is where claims die: submit, then wait 4–8 weeks; read a rejection that contradicts the airline's own operations log; challenge it; wait another 4 weeks; escalate — but only when the escalation is ripe.

The work is small. The calendar is enormous. That's the persistence gap: whether you get what the policy says you're owed depends less on the merits than on whether you have the knowledge, time and stubbornness to keep coming back for three months.

The thesis The hard part of agentic AI isn't making an agent reason once. It's making it reliably stay with a real-world problem until the problem is actually resolved.

Agents reason. Workflows remember.

A stateless caseworker (4 tools, 10-step loop) and an independent verifier do the thinking; a Restate workflow owns the case. The agent forgets everything when it returns; the workflow survives kill -9, waits days for a human, and cannot execute a consequential action twice.

That split is why NanoClaw was evaluated and rejected — its persistent sessions would have competed with the workflow for owning case memory.

Tools, not an oracle. No tool returns a verdict. A check_eligibility() backed by a rules engine would have made the benchmark a test of whether the agent can call one function.

Independent verification. The verifier sees the case and the policy but not the caseworker's working — a second opinion, not a reviewer inheriting a wrong turn. It can reject; a rejection citing no clause is downgraded to a pass.

Human approval that is enforced, not requested. Every consequential action blocks on a durable promise, and the rejection branch returns before any outbound call is reachable.

Results — two suites Suite A · reasoning · 28 synthetic cases.

Primary metric is Case Resolution Accuracy: a case counts only when the next action, the compensation figure and the other entitlements are all correct at once.

System CRA Model calls Cost Baseline — one direct prompt, no tools, no verifier 0.82 28 $1.37 Full agent — tools, loop, independent verifier 0.93 102 $3.77 Same model, same policy, same cases, same schema, same grader, same endpoint.

The best possible constant answer on this suite is 0.25. Read +0.11 as directional, not significant — it is three cases, and this project has no valid variance estimate; the one it had was withdrawn in FAILURES.md F-008.

Suite B · durability · 6 failure-injection scenarios.

6/6 passed (6/6 again inside Docker), workflow completion 1.00, failure recovery 1.00, state preserved after kill -9 1.00, and 0 duplicate consequential actions.

The load-bearing one is D06: SIGKILL the worker inside the submission window — the carrier received 2 submission attempts and 1 claim landed, because the idempotency key is generated inside a durable step and is stable across replay.

In D05, when the human refuses, the carrier is called zero times, not once and then reversed.

The baseline scores nothing here, and that is not reported as a win.

The biggest improvement On the reasoning axis, the independent verifier — not the tools.

The agent made 1.4 tool calls per case and called calculate three times in 28 cases. The verifier is isolated by case R16, a partial settlement that fails under the baseline and under tools-only and passes only with verification.

It costs 3.6× the model calls for +0.11.

But the improvement that mattered wasn't a component. It was answering where does the case remember itself? and being strict about it — which is why there is no database, why FastAPI holds no case state, and why a crash produces two attempts and one claim instead of two claims.

The main failure mode we did not fix The agent asserts money it never computed.

Both remaining wrong cases fail on duty-of-care arithmetic; calculate was available on both and called on neither, despite a bolded prompt instruction with an example.

It does the sums in its head because it is capable of doing them in its head and confident about it.

That is invisible in an aggregate score and it is the one that reaches the passenger: a confident, well-cited, correctly-formatted claim letter with the wrong number in it.

Corrective experiment EXP-005 is specified and recorded as unrun.

Hot take When a control moves and you did not touch the control, stop theorising about the treatment.

Three times this project reported infrastructure as capability: a required schema field zeroed a well-formed verdict; six cases were never sent to the model and were scored as six wrong answers, standing as the headline comparison for two days; a gateway served a different model while every result file recorded the one that was requested.

All three looked exactly like reasoning results. What caught all three was the baseline moving when nothing about the baseline had changed.

A missing answer is not a wrong answer. If your grader can't tell them apart, it will report your infrastructure as your agent's reasoning.

The model field is not provenance. The endpoint is.

Re-run your control, not just your treatment.

And the thing I got wrong at the start: I assumed the reasoning would be hard and durability would be plumbing.

A single prompt already handles most of these cases. What a single prompt cannot do is still be holding the case in week six.

Reproducing it git clone https://github.com/yablokolabs/c2c-agent.git && cd c2c-agent

make setup

make test — 165 tests, no model calls, no services

make reproduce — baseline + agent + comparison

The headline result needs no Restate server, no Docker and no Telegram.

For the durability half:

make up && make failure-tests && make down

Deliverables Deliverable Location 01 · Code + improvement changelog the repo · IMPROVEMENT_CHANGELOG.md 02 · Reproduction guide REPRODUCTION_GUIDE.md 03 · Solution video script in docs/DEMO_SCRIPT.md, submitted separately 04 · Agent trajectories trajectories/README.md Failure journal: FAILURES.md

Experiments including removed ones: experiments/

Raw results: evaluation/results/

Raw results are never overwritten, each carrying its commit, model, backend and prompt digests.

Built With

Share this project:

Updates

Submission history