Inspiration

The idea isn't ours. It came from Ruchi, a friend studying law at {NLU}, who kept describing the same problem from the other side of it. The project is named after her.

She was deep in moot court season — written submissions, memorials, rounds of edits. And the thing that kept going wrong was never a bad argument. It was citations. A case cited for a proposition slightly wider than what the court actually held. A paragraph number that pointed at the wrong paragraph. A line quoted as the ratio that turned out to be counsel's submission, recorded by the judge and then rejected two paragraphs later. She had a word for it — adding masala on top of what the court said.

What made it stick was that none of this is dishonesty. Verifying one citation properly means opening a 300-page judgment and reading it. Nobody does that at the rate briefs get written, and law students least of all. Then we read about a petition in the Delhi High Court that relied on paragraphs 73 and 74 of a judgment that has 27 paragraphs, and about courts beginning to sanction advocates for citing cases that were never decided at all.

And general-purpose AI makes it worse, not better. Ask a chatbot for authority and you get a citation in the right format, with plausible party names, for a case that does not exist.

So: build the thing that reads the 300 pages. Ruchi told us what the twelve ways a citation lies actually are — we just wrote them down in code.

## What it does

Ruchi takes a brief and asks three questions separately about every citation in it:

  1. Does the case exist?
  2. Which paragraph is actually being relied on?
  3. Does that paragraph support the proposition to the extent claimed?

That third one is the whole point. Most citations in a bad brief are real cases — they're just stretched.

It runs the same machinery in reverse: give it a proposition and it finds the judgment that backs it, and the line. Run it again with the sign flipped and it finds the judgment that says the opposite — what opposing counsel will cite. And there's a drafting gate: each sentence you intend to argue is bound to an authority or refused outright, so nothing reaches the document that the verifier couldn't confirm.

On top of all that sits an agent, so you don't have to know which check you want. You ask "the other side cited this for that proposition — is any of it true?" and it works out that this is four checks and an ordering.

The corpus is the entire Supreme Court of India: 38,005 judgments with full text, 707,647 paragraphs, 1950 to 2025, from AWS Open Data under CC-BY-4.0.

## How we built it

The engine came first and the agent came last, which turned out to be the right order.

The engine is a LangGraph state graph — eight nodes in a fixed sequence: resolve the citation, check whether it's still good law, load the text, locate the paragraph, compare the scope of the claim against the holding, work out whose voice the passage is in, check applicability, assemble a verdict. Around it: a BM25 index over all 707,647 paragraphs, a citator built from citation edges, and a quote verifier.

Some pieces we're glad we did the hard way:

  • The citator applies a bench-strength rule, so a smaller bench can never be reported as having overruled a larger one. It's arithmetic, not a model.
  • Retrieval drops any passage that isn't the court speaking before ranking. An advocate states a rule more baldly than a judge ever will, so counsel's submission out-matches the actual holding on keywords. Offering it as authority would hand you the exact mistake the verifier exists to catch.
  • Every model-produced sentence is string-matched against the stored judgment. If the match fails, the claim is downgraded, however confident the model sounded.

The agent layer is the Strands Agents SDK: one Agent, eight @tool functions, and a system prompt. Each tool opens a database session, calls one function that already existed, and returns JSON. Bedrock is the default model provider, with the project's existing provider chain behind it through LiteLLM. When no AWS credential resolves it falls back — and reports which model answered, in every response. That last detail turned out to matter far more than we expected.

The rule we held to: the agent chooses the question, never the answer. The verification graph still runs its eight nodes in the same fixed order. The agent can't reorder one, skip one, or overrule a verdict, and it has no path to the corpus except through a tool.

## Challenges we ran into

Counsel's voice vs. the court's holding. Our first retrieval kept returning paragraph 2 of judgments when the answer was paragraph 3 — because paragraph 2 was the advocate's submission, stated without qualifications, and therefore a better keyword match. This is the single hardest thing in the project and it needed a rhetorical-role pass over every paragraph.

The metadata lies about bench strength. The open-data judge column names only the presiding judge, so counting from it under-reports the bench. The coram printed above the headnote is the real thing, which meant parsing the PDFs properly.

The reporters' own words are in the text. Law reports open with an editorial headnote that isn't the court speaking. Treating it as judgment text would mean verifying quotes against a publisher's summary.

Free-tier quotas are brutal. We spent most of the build on 20 model requests per day. We exhausted it testing the deployed agent and watched our own live demo start returning errors, which is a memorable way to learn what your dependency actually is.

A sampling preference took the whole agent down. We pinned temperature to zero for every model, because the agent's job is choosing which of eight checks to run and two identical questions shouldn't produce two different audits. Then we pointed the deployed chain at a GPT-5 model, which accepts only temperature=1 — and the library raises an error rather than dropping the parameter. Every single question came back as a stack trace instead of an answer. We found it on deadline day, by asking the deployed agent a question rather than trusting that a green test suite meant a working system.

Deploying it taught us things nobody warns you about. Our container couldn't reach the EC2 instance role because the IMDS hop limit defaults to 1. Our compose file passes an explicit allowlist of environment variables, so a setting not named in it silently never arrives — and the agent falls back to a different model and keeps answering, which is the failure that looks like success.

## Accomplishments that we're proud of

That it refuses. We asked the deployed system to find authority for "a notice under Section 106 is mandatory before a suit for eviction." It chained three checks, searched 707,647 paragraphs, found a passage containing those exact words — and then declined to put it behind the proposition, because the judgment merely recorded as a fact that notice had been served in that case. It does not hold that notice is mandatory.

That citation would have looked perfect in a brief and collapsed in court.

Eight of the twelve failure modes are decided with no language model at all — parsing, indexing and arithmetic.

Three states, never collapsed: supported, checked-and-not-supported, and not checked. A treatment of "good law" carries unchecked: true when nothing in the corpus has ever cited that judgment — good law by default, not by evidence. A contrary lead nobody has read carries confirmed: null, and null is not false. Each of these is asserted by a test, because an edit that collapsed one would otherwise leave the suite green.

## What we learned

That putting an agent in front of a citation checker risks the exact failure the checker exists to catch.

A language model asked whether an Indian judgment is still good law already believes it knows. The belief is fluent, confident, and worth nothing. Most of our work on the agent wasn't wiring tools up — that part took an afternoon. It was writing a prompt and a set of return shapes that make answering from memory harder than calling the tool.

We also learned to distrust "it works", twice over. Once when our deployed agent answered beautifully while running on a completely different model than we believed, because a fallback had fired silently. And again when a one-line sampling preference turned every deployed request into an error while every test on our machines still passed. Both were invisible from where we were standing. Now every answer reports which model produced it, and we check the deployed thing by asking it questions, not by reading our own test output.

And the thing Ruchi taught us, which we didn't understand at the start: in law, the useful answer is often a refusal. "Nothing in this corpus supports that sentence" is a result a student can act on. A confident wrong citation is not.

## What's next for Ruchi

  • Bedrock AgentCore for a managed runtime. The provider is already Bedrock and the tools are already plain functions, so this is deployment, not a rewrite.
  • A session manager, so a follow-up question doesn't have to restate what it refers to.
  • High Court coverage. Right now the corpus is the Supreme Court only, and the agent says so rather than guessing past it.
  • The evaluation at scale. We have a held-out set of 313 items and, for most of the build, a free tier that allowed 20 requests a day. The pipeline is proven; the measurement isn't, and we'd like to publish real numbers rather than claims.
  • Getting it in front of actual moot court teams — starting with Ruchi's.

Built With

Share this project:

Updates

Submission history