Inspiration
I was ranking online contests by prize divided by entrants, looking for the best place to spend a month of work. One topped the list by a factor of five. It advertised $149,525 in prizes and its machine readable listing declared five cash prizes, with about five hundred people registered.
Under every single prize tier, fifteen times on the same page, it says
⚠️ THIS IS NOT A CASH PRIZE. This award consists of sponsor-provided product credits, subscriptions, software licenses, domains, or other non-cash benefits and cannot be exchanged or redeemed for cash.
The real cash was zero. Nothing was hidden. The number I was about to act on had been typed in by the party who benefits from it being large, and the page contradicted it in plain sight, one scroll below.
Then I read the long legal rules of twenty two contests by hand and concluded
that twelve of them were open to an adult in my country. The first working
build of this agent came back with Students only on pages I had cleared
myself three hours earlier. The real number was eight. I had read the
longest document on each page and missed the shortest field, four words in the
eligibility card beside the register button. The tool's first useful act was
to correct its own author.
What it does
Asterisk reads an offer the way a careful person would if they had two hours, and reports two different things.
Contradictions. The page announces something loudly and takes it back quietly, somewhere else on the same page. Free that becomes charged. Unlimited that becomes capped. Cash that is not cash.
Quiet conditions. Nothing on the page promises otherwise, and a clause binds you anyway. Reserved to students. Automatic renewal. Non refundable. Rights on what you submit.
Every line it prints is a verbatim quote from the page. A finding that cannot be pointed at is dropped before you see it.
Who it is for
Anyone about to click a button that costs them something and does not have two hours. Practically, that is a person comparing subscriptions, a person about to enter a competition, a person accepting a free trial that wants a card first.
What using it is actually like
One command, and the deterministic half needs no key and no install.
python cli.py --offline https://example.com/offer
It prints verbatim quotes with their place on the page. --html writes a self
contained report you can open or forward to someone who will never install
anything. --json makes it scriptable, and the exit code is part of the
interface, 0 for nothing serious, 1 when something serious was found, 2 when
the page could not be read.
That last code is the design decision I care most about. Two of the ten
consumer pages I tried draw their prices in the browser, so a plain fetch sees
an empty shell. A tool that prints nothing found there is worse than no tool,
because silence and a clean bill of health look identical to a reader. It exits
2, says why, and offers --browser.
--watch stores a snapshot and reports what changed since the last one, which
is how you catch an offer that quietly grows a clause.
And you do not have to take any of this on trust. Eight entries on real pages are published, seven of them full reports and one a refusal. Consumer offers, a music subscription, a money transfer service, storage plans, a commerce price table, and contest rule pages. The eighth is there because the tool REFUSED to audit it, so the gallery shows the failure mode next to the successes. The builder rebuilds them from the live pages and fails if one is missing.
https://thibaudlepan77-svg.github.io/asterisk/
How I built it
Four stages, and the two most interesting were rewritten because a measurement said the first design was wrong.
The page becomes prominence scored lines, then sentences. Deterministic rules catch the family of contradiction that consumer law forces into fixed wording, free and instant and inventing nothing. A model layer catches everything the formulas miss, because most fine print was never standardised.
The agent, built on Strands Agents SDK, exists for the part that needs
judgement. Three tools, a system prompt that hands it the deterministic audit
rather than asking it to run one, and a BeforeToolCall hook that enforces the
page budget instead of requesting it. A clean offer page means nothing on its own, the clause that costs
you is usually behind a small grey link. The agent ranks the links that could
hold binding conditions, follows the ones worth a request inside a page
budget, and stops when another page would add nothing.
Challenges
The first design compared a loud block with a quiet block. It scored zero on live pages, because the denial usually sits inside the promise block, two sentences later. The unit that carries a contradiction is the sentence.
Modern markup shatters a sentence. The currency symbol, the digits and the
words in cash arrive as three sibling elements, and a matcher that keeps them
apart finds nothing at all. Fragments are merged back into lines first.
The model skipped the page it was asked to audit. Told to audit the starting page and then explore, it went straight to the linked pages, found them clean, and called a page carrying three contradictions clean. The starting audit is now deterministic and handed to the model. What can be decided without judgement should not be left to judgement.
The page budget was a sentence in a prompt. The task text said Budget, at
most 4 pages in total and nothing checked it. A model that finds the fourth
page interesting fetches a fifth. It now lives in a Strands BeforeToolCall
hook, where cancel_tool turns a refusal into a refusal, and the decision sits
in twenty lines that need no model to test. Writing the hook turned up the hole
it did not cover, verify_quote fetched any page it had not seen, so one more
page was always one tool call away. A limit one door enforces and another
ignores is not a limit.
A tool that lets a model attest to its own output has verified nothing.
The agent had a verify_quote tool. It called it. It passed. Its answer still
wrote licences where the page says licenses, and organisations where the
page says organizations, with every line labelled verified. The model had
checked the real string and then written a tidied one, because tidying prose is
what a language model does. Verification moved to after the last token, on
the text the reader sees. An answer containing an unsupported quote exits non
zero.
Accomplishments
It is measured, and the measurements are in the repository and rerunnable.
Four labels are scored, on two sets, and every reference label is read from a source the auditor never parses.
On 22 contest pages the rules were written against, students only,
prize is not cash and team required all score precision 1.00 and recall
1.00, and a video is required scores 0.94 and 1.00.
On 40 pages drawn at random from 13,632 finished contests that the rules were
never written against, students only holds 1.00 and 1.00 on four positives,
team required holds 1.00 and 1.00 on five, a video is required scores 0.88
and 1.00 on seven, and the non cash detector raises zero false alarms on all
forty.
The two labels above the line were added because the tool prints eleven families of restriction and two of them had ever been scored. Nine tenths of what it said carried no number in front of it, and a number nobody has is indistinguishable from a number that is bad.
The first version of the video reference was wrong and the tool was right. It scored the tool at 0.27 precision, and every one of its eleven false alarms was a page that plainly asks for a video. The reference had been built from the one wording its author had read, which is the mistake this project keeps finding in its own rules. The one false alarm still left on each set is the reference too. Both quotes are in the README and can be checked in ten seconds. The reference was not rewritten a third time on purpose, because tuning a ground truth until the tool looks perfect is fitting the answer.
Taken off its home ground onto ten ordinary consumer offers, it quoted a music
service back at itself with Try 3 months of Premium Individual for $0, then
$12.99/month, and an online bank advertising no fees with A fee of 2% (or
minimum fee £1) applies per transaction after your rolling monthly limit.
What I learned
Two of those ten pages draw their prices in the browser, so a plain fetch sees
a one line shell. Reporting nothing found on those would be the worst thing
this tool could do, because silence and a clean bill of health look identical
to a reader. It now refuses to be silent about it, exits with a distinct code,
and offers a rendered fetch.
What is next
Following a condition trail across a login, and a second language. The deterministic rules are English first, and the formulas they match are the English ones.
Built with
Python, Strands Agents SDK, an OpenAI compatible endpoint, Playwright for the optional rendered fetch. No other dependency on the deterministic path.
Built With
- amazon-web-services
- openai
- playwright
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.