Inspiration

Legal and procurement teams negotiate the same clauses with the same kinds of vendors again and again, and what they learn stays in old redlines, emails and senior colleagues' heads. When a lawyer leaves, the knowledge leaves too. We wanted an agent that doesn't just draft a redline but remembers what actually closed: which fallback a vendor accepted, which ask always gets rejected, and when a vendor's behavior changed. [FILL IN: one or two sentences of your own motivation, such as a personal experience with contracts, procurement or repeated paperwork.]

What it does

Precedent is a negotiation copilot for a buyer-side legal team. It negotiates a six-clause SaaS agreement (liability cap, indemnity, data protection, payment terms, termination, audit rights) against a simulated vendor. You pick a vendor and a memory mode (none, recent log, or Hindsight), then:

  1. In Hindsight mode, the agent reflects on past outcomes with that vendor type and opens with a written briefing.
  2. For each clause it recalls relevant memories, shown in the UI as citation chips (M1, M2, and so on), and proposes a position with a cited rationale.
  3. A human lawyer approves, edits or regenerates the proposal (or auto-approve is on).
  4. The vendor counters according to hidden rules the agent never sees.
  5. Agreed outcomes are retained to Hindsight, the deal scoreboard updates live, and the Memory page consolidates beliefs with proof counts across runs.

The Experiments page replays the same negotiation sequence under all three memory conditions and charts rounds, value captured and prediction error, which is the measurable case that memory helps.

How we built it

  • Memory: Hindsight (Cloud) for retain, recall and reflect. We use its observations to consolidate patterns across deals, timestamps so changes over time are visible, and directives such as "never present output as legal advice."
  • Agent: Groq-hosted models, openai/gpt-oss-120b as primary and qwen/qwen3.8-27b as fallback, with a forced decide_action tool call, schema validation and retries.
  • Backend: FastAPI with async SQLAlchemy and SQLite, 19 REST and SSE endpoints, streaming agent events to the browser live.
  • Frontend: React 18, TypeScript (strict) and Tailwind with a custom component library. The centerpiece is a redline view with struck-through and inserted wording next to a memory rail that works like margin notes.

Proving that memory helps

Vendors are simulated, so the evaluation has to be honest. Each vendor type has hidden rules, and the agent must discover them by negotiating. Each clause has four positions, level 1 being best for us. If the vendor's true minimum acceptable level is $m_c$ and we close at level $\ell_c$, a clause with weight $w_c$ is worth

$$v_c = w_c \cdot \frac{4-\ell_c}{3}$$

and the share of available value captured in a negotiation is

$$\text{captured} = \frac{\sum_c v_c}{\sum_c w_c\,\frac{4-m_c}{3}} \times 100\%.$$

We also track prediction error, $|\hat{m}_c - m_c|$, and the rounds needed per clause. The same vendor sequence runs under three conditions: no memory, a naive log of recent events stuffed into the prompt, and Hindsight. Partway through, one vendor type changes a rule, so we can see whether memory adapts.

Challenges we ran into

  • Keeping the evaluation honest. The hidden vendor rules live in one module, are read only by the simulator and an explicit reveal endpoint, and are scrubbed from anything retained to memory. Live mode shows "Deal score: X of 100" instead of a percentage of the hidden maximum, so it can't reveal the rules either.
  • Retained data is asynchronous in principle. We designed a settle mechanism with three fallbacks, then measured the real behavior: on Hindsight Cloud, retain is synchronous and content is recallable about half a second later. We recorded the measured latencies rather than assuming.
  • Unreliable model availability. Groq's free tier has a daily token cap, so long experiments can hit it. Retries and the fallback model absorbed this. We also built a test where the LLM is forced offline, and the run completes with visible degraded markers instead of failing silently.
  • A known limitation. Toggling auto-approve mid-run doesn't release a clause already waiting at the approval gate, because the flag is read when a run starts.
  • [FILL IN: other real problems you hit and how you fixed them.]

Accomplishments that we're proud of

  • Built and verified end to end against real Hindsight Cloud and Groq: a live smoke test passing 8 of 8 checks, live negotiations, a completed three-condition experiment, and the full UI flow tested in the browser.
  • Zero leaks of the hidden rules. An adversarial recall probe found none, and retained text is scrubbed automatically.
  • Zero degraded decisions in a full experiment, even though 29 calls needed the fallback model.
  • A tested quality floor: 27 backend tests, 11 frontend tests and a strict TypeScript check.
  • Memory you can see. Every proposal cites specific past deals, and the memory overlay on each clause shows what Precedent expects the vendor to accept.
  • Results: [FILL IN from your experiment: for each of the three conditions, the average rounds per clause, value captured (%) and prediction error over the last 8 negotiations, plus how many negotiations Hindsight needed to recover after the rule change. Report the real numbers, including anywhere Hindsight did not win.]

What we learned

  • Memory is only convincing if it is measurable. A before-and-after chart beats any feature.
  • Baselines matter. A "recent history in the prompt" baseline is a real competitor, so we compare against it, not only against no memory.
  • Measure infrastructure behavior instead of assuming it, such as retain timing and rate limits.
  • Honesty is a design feature. Marking degraded decisions and hiding ground truth made the results more credible.
  • [FILL IN: what you learned about Hindsight itself, for example how observations and reflect behaved in practice.]

What's next for Precedent

Importing real redline histories, splitting uploaded contracts into clauses automatically, sharing memory across a team so a junior lawyer benefits from a senior's past deals, and modeling trade-offs between clauses.

Limitations: Vendors are simulated, so results show that the mechanism works, not that it will work on real negotiations. Precedent is decision support for lawyers. It never replaces their judgment and never gives legal advice.

Built With

Share this project:

Updates

Submission history