Inspiration

I was ranking online contests by prize divided by entrants, looking for the best place to spend a month of work. One topped the list by a factor of five. It advertised $149,525 in prizes and its machine readable listing declared five cash prizes, with about five hundred people registered.

Under every single prize tier, fifteen times on the same page, it says

⚠️ THIS IS NOT A CASH PRIZE. This award consists of sponsor-provided product credits, subscriptions, software licenses, domains, or other non-cash benefits and cannot be exchanged or redeemed for cash.

The real cash was zero. Nothing was hidden. The number I was about to act on had been typed in by the party who benefits from it being large, and the page contradicted it in plain sight, one scroll below.

Then I read the long legal rules of twenty two contests by hand and concluded that twelve of them were open to an adult in my country. The first working build of this agent came back with Students only on pages I had cleared myself three hours earlier. The real number was eight. I had read the longest document on each page and missed the shortest field, four words in the eligibility card beside the register button. The tool's first useful act was to correct its own author.

What it does

Asterisk reads an offer the way a careful person would if they had two hours, and reports two different things.

Contradictions. The page announces something loudly and takes it back quietly, somewhere else on the same page. Free that becomes charged. Unlimited that becomes capped. Cash that is not cash.

Quiet conditions. Nothing on the page promises otherwise, and a clause binds you anyway. Reserved to students. Automatic renewal. Non refundable. Rights on what you submit.

Every line it prints is a verbatim quote from the page. A finding that cannot be pointed at is dropped before you see it.

Who it is for

Anyone about to click a button that costs them something and does not have two hours. Practically, that is a person comparing subscriptions, a person about to enter a competition, a person accepting a free trial that wants a card first.

What using it is actually like

One command, and the deterministic half needs no key and no install.

python cli.py --offline https://example.com/offer

It prints verbatim quotes with their place on the page. --html writes a self contained report you can open or forward to someone who will never install anything. --json makes it scriptable, and the exit code is part of the interface, 0 for nothing serious, 1 when something serious was found, 2 when the page could not be read.

That last code is the design decision I care most about. Two of the ten consumer pages I tried draw their prices in the browser, so a plain fetch sees an empty shell. A tool that prints nothing found there is worse than no tool, because silence and a clean bill of health look identical to a reader. It exits 2, says why, and offers --browser.

--watch stores a snapshot and reports what changed since the last one, which is how you catch an offer that quietly grows a clause.

And you do not have to take any of this on trust. Eight entries on real pages are published, seven of them full reports and one a refusal. Consumer offers, a music subscription, a money transfer service, storage plans, a commerce price table, and contest rule pages. The eighth is there because the tool REFUSED to audit it, so the gallery shows the failure mode next to the successes. The builder rebuilds them from the live pages and fails if one is missing.

https://thibaudlepan77-svg.github.io/asterisk/

How I built it

Four stages, and the two most interesting were rewritten because a measurement said the first design was wrong.

The page becomes prominence scored lines, then sentences. Deterministic rules catch the family of contradiction that consumer law forces into fixed wording, free and instant and inventing nothing. A model layer catches everything the formulas miss, because most fine print was never standardised.

The agent, built on Strands Agents SDK, exists for the part that needs judgement. Three tools, a system prompt that hands it the deterministic audit rather than asking it to run one, and a BeforeToolCall hook that enforces the page budget instead of requesting it. A clean offer page means nothing on its own, the clause that costs you is usually behind a small grey link. The agent ranks the links that could hold binding conditions, follows the ones worth a request inside a page budget, and stops when another page would add nothing.

Challenges

The first design compared a loud block with a quiet block. It scored zero on live pages, because the denial usually sits inside the promise block, two sentences later. The unit that carries a contradiction is the sentence.

Modern markup shatters a sentence. The currency symbol, the digits and the words in cash arrive as three sibling elements, and a matcher that keeps them apart finds nothing at all. Fragments are merged back into lines first.

The model skipped the page it was asked to audit. Told to audit the starting page and then explore, it went straight to the linked pages, found them clean, and called a page carrying three contradictions clean. The starting audit is now deterministic and handed to the model. What can be decided without judgement should not be left to judgement.

The page budget was a sentence in a prompt. The task text said Budget, at most 4 pages in total and nothing checked it. A model that finds the fourth page interesting fetches a fifth. It now lives in a Strands BeforeToolCall hook, where cancel_tool turns a refusal into a refusal, and the decision sits in twenty lines that need no model to test. Writing the hook turned up the hole it did not cover, verify_quote fetched any page it had not seen, so one more page was always one tool call away. A limit one door enforces and another ignores is not a limit.

A tool that lets a model attest to its own output has verified nothing. The agent had a verify_quote tool. It called it. It passed. Its answer still wrote licences where the page says licenses, and organisations where the page says organizations, with every line labelled verified. The model had checked the real string and then written a tidied one, because tidying prose is what a language model does. Verification moved to after the last token, on the text the reader sees. An answer containing an unsupported quote exits non zero.

Accomplishments

It is measured, and the measurements are in the repository and rerunnable.

Four labels are scored, on two sets, and every reference label is read from a source the auditor never parses.

On 22 contest pages the rules were written against, students only, prize is not cash and team required all score precision 1.00 and recall 1.00, and a video is required scores 0.94 and 1.00.

On 40 pages drawn at random from 13,632 finished contests that the rules were never written against, students only holds 1.00 and 1.00 on four positives, team required holds 1.00 and 1.00 on five, a video is required scores 0.88 and 1.00 on seven, and the non cash detector raises zero false alarms on all forty.

The two labels above the line were added because the tool prints eleven families of restriction and two of them had ever been scored. Nine tenths of what it said carried no number in front of it, and a number nobody has is indistinguishable from a number that is bad.

The first version of the video reference was wrong and the tool was right. It scored the tool at 0.27 precision, and every one of its eleven false alarms was a page that plainly asks for a video. The reference had been built from the one wording its author had read, which is the mistake this project keeps finding in its own rules. The one false alarm still left on each set is the reference too. Both quotes are in the README and can be checked in ten seconds. The reference was not rewritten a third time on purpose, because tuning a ground truth until the tool looks perfect is fitting the answer.

Taken off its home ground onto ten ordinary consumer offers, it quoted a music service back at itself with Try 3 months of Premium Individual for $0, then $12.99/month, and an online bank advertising no fees with A fee of 2% (or minimum fee £1) applies per transaction after your rolling monthly limit.

What I learned

Two of those ten pages draw their prices in the browser, so a plain fetch sees a one line shell. Reporting nothing found on those would be the worst thing this tool could do, because silence and a clean bill of health look identical to a reader. It now refuses to be silent about it, exits with a distinct code, and offers a rendered fetch.

What is next

Following a condition trail across a login, and a second language. The deterministic rules are English first, and the formulas they match are the English ones.

Built with

Python, Strands Agents SDK, an OpenAI compatible endpoint, Playwright for the optional rendered fetch. No other dependency on the deterministic path.

Built With

Share this project:

Updates

Submission history