Inspiration

We went looking for funding for a small project. In one afternoon we read the rules of three open calls advertising $224,000 between them. Here is what the rules actually said:

Advertised The sentence that ended it
$148,445 "Grand Champion: $300 in Featherless AI credits" — the prize is not money
$35,320 "Prizes: TBD. Further announcements will be made soon!" — and students only
$40,000 "At least one team member must attend the NeurIPS 2026 presentation in person."

Three for three. Each one findable in ten minutes of reading, and each one capable of eating a week of work if you don't do that reading.

Then we thought about who this actually happens to. Not us — we had an afternoon to spare. A neighbourhood library. A food bank. An all-volunteer group with two people and no grants officer. They write the whole application. They find out at the end, if they find out at all.

Clausewitz is that ten minutes of reading, done for every call at once, with the disqualifying sentence quoted back to you.

What it does

You give it your organisation's real profile — country, legal form, whether it can send someone somewhere in person, whether it can pay costs up front and be reimbursed later, whether a prize paid in credits is any use to it. Then you give it open calls.

It sorts them into three buckets, and every exclusion cites the text that caused it:

Screening 8 open calls for: Biblioteca Vecinal San Andres

  eligible 3   not eligible 3   needs a human 2

NOT ELIGIBLE (3)   -- with the clause, not a label
   Global Innovation Build Challenge V2
        "Grand Champion: $300 in Featherless AI credits"
        -> the award is not money, and this profile needs cash rather than credit

   AWS Trainium Frontier Competition
        "At least one team member must attend the NeurIPS 2026 presentation in
         person."
        -> attendance in person is required and this profile cannot travel

CANNOT DECIDE (2)   -- and saying so is the point
   Fondo Municipal de Cultura 2026
        no eligibility criteria could be extracted at all

3 applications not worth writing. At roughly six hours each, about 18 hours back.

How I built it

The model reads. The code decides. Everything follows from that split.

A language model is genuinely good at finding the eligibility sentence buried on page four between the sponsor logos and the schedule. It is genuinely bad at being trusted with the answer, because when it is wrong it is wrong fluently.

So the Strands agent gets exactly one job and one tool. It reports each demand the call makes, quoting the sentence it came from. It never returns a verdict.

Then every quote is checked against the source text, character for character. A requirement the model invented quotes nothing that exists, fails the check, and is dropped. A hallucination cannot become a rejection. At worst it becomes a missing data point, which pushes the call towards a human.

That guarantee is not a claim about the prompt. It is a property of the code, and it is tested 61 times without ever calling a model. --audit makes it runnable: it prints each quote against its source and exits non-zero if any is not found.

The screening layer underneath has no third-party dependencies at all — no model, no network, no credentials. Clone it and run it.

Three buckets, because two is a lie

Most screeners answer yes or no. The third bucket is the point.

  • ELIGIBLE — every rule that was read, passed
  • EXCLUDED — a rule blocked, and here is the sentence
  • UNDECIDABLE — something could not be read, resolved, or understood

A call that says nothing about who may enter is not a call that admits everyone. It is one that needs a person.

Every ambiguity fails towards UNDECIDABLE, never towards EXCLUDED. Wrongly telling someone not to apply costs them a grant. Wrongly asking a human to look costs them a minute.

Challenges I ran into

All four of these were found by running the thing. None was visible by reading it, and every one failed silently.

1. It told a library not to apply to a fund that wanted exactly that library. A call accepting "registered nonprofit organisations" was compared, as an exact string, against a profile whose legal form is nonprofit. Prose never matches a controlled vocabulary. The verdict was EXCLUDED, and it was confidently wrong in the one direction that costs a user real money.

2. Country lists were iterated as characters. The model returned "Brazil, Italy, Quebec" as a string. The check did for c in value. Python iterated the letters, so the comparison set became {'B','R','A','Z'...} and a two-letter country code could never match. A call that excludes you would have cleared you, with nothing logged anywhere.

3. The same bug again, somewhere else. I had fixed the instance, not the class. The legal-form check had the identical line, and the demo caught it by printing the call requires ['a','d','i','l','n','u','v']. So the regression test does not test the two known sites — it sweeps every check that takes a list and requires a string and a list to give the same answer. The third occurrence cannot be written now without a test failing.

4. A network failure looked exactly like a vaguely written call. Both produce zero requirements, and zero requirements printed as "no eligibility criteria could be extracted at all" — a true sentence about a document, a false one about a 503. Now they are different outcomes with different exit codes.

There was also a reproducibility trap worth naming: gemini-2.5-flash returns 404 "no longer available to new users" on an API key issued today, while still working for anyone who onboarded earlier. Pinning it would have broken for every judge and for nobody on this side.

Accomplishments that I'm proud of

Quebec does not become Canada. Excluding Quebec does not exclude Canada, so a sub-national jurisdiction that cannot be resolved comes back as a question rather than being mapped to the nearest country. Half of contest rules carve out Quebec — including this hackathon's own eligibility list, which is where I first saw it.

More generally: the tool refuses to guess. Every place it cannot resolve, every legal form it does not recognise, every requirement kind it has no rule for, comes back named, in the open, addressed to a person.

What I learned

That the dangerous failure is the quiet one. Bug 3 was loud — it printed a list of single letters and anyone would have caught it. Bug 2 was the same line of code, pointing the other way, and it produced a clean, confident, wrong "you are eligible". Same root cause. Only one was survivable.

And that a screener's failure direction is a design decision, not an accident. Once you say out loud which wrong answer costs the user more, the whole architecture follows.

What's next for Clausewitz

  • More rule kinds. Every this screener has no rule for: line in the output is a feature request the tool wrote for itself.
  • A shared profile format, so a small organisation fills it in once and screens forever.
  • Watching calls over time. The genuinely interesting event is a call whose rules change after you started writing.

Disclosure of pre-existing work

The rules require that projects be newly created during the submission period, and that pre-existing work be disclosed. Stating it plainly:

  • All code in this repository was written during the submission period. No file was carried in from an earlier project.
  • The ideas were not invented here. I previously built an unrelated tool (a phone-quote agent) that established two patterns reused as design knowledge: printing exclusions rather than silently filtering them, and validating narrow fields against their expected shape. No code, tests or text were copied.
  • The three real calls quoted in fixtures/calls.json are calls I read while looking for funding. The wording is theirs; the screening is mine.
  • Development used AI coding assistants, which the rules permit explicitly.

Built With

Share this project:

Updates

Submission history