ProofMark project overview
Project name: ProofMark
Elevator pitch: ProofMark turns AI-generated Singapore legal answers into auditable claims, checks their citations and exact evidence, and sends unsupported or uncertain findings to lawyer review instead of asking another black-box model to be the final judge.
Thumbnail: proofmark-project-thumbnail.png
Inspiration
Legal AI does not fail only by inventing cases. The harder failures often look completely credible: a real case may be paired with the wrong name, an accurate quotation may be taken from a party's submission rather than the court's reasoning, a genuine authority may not support the proposition attributed to it, or an answer may omit an important qualification altogether.
Most citation checkers stop after confirming that an authority exists. At the other extreme, asking a second general-purpose language model to judge the first simply moves the trust problem. We wanted a system that could show its work, use deterministic checks wherever possible, and be explicit about where legal judgment still belongs to a human reviewer.
That led us to ProofMark: an implementation of our VERITAS SG assurance framework, scoped to Singapore employment restraint-of-trade law.
What it does
ProofMark evaluates an AI-generated legal answer through four practical questions:
- Does the authority exist and is it correctly identified?
- Does the cited material actually support the claim?
- Does it carry the legal meaning and weight claimed?
- What may the answer have missed?
A user pastes an AI-generated answer and can also provide the original question and facts. ProofMark then breaks the answer into atomic claims, maps each claim to its cited authorities, and checks citation identity, metadata, pinpoints, quotations, proposition fit, modality, context, currency, and bounded omission signals.
The dashboard returns an evidence-linked review rather than a single opaque score. Each finding shows the exact stored passage, source provenance, decision rule, confidence boundary, and recommended lawyer action. Unknown citations remain unverified unless an official negative check supports a stronger conclusion. Potentially relevant uncited authorities are presented only as research leads requiring source and treatment review; they never change a verdict automatically.
How we built it
We built a full-stack, locally runnable MVP with a FastAPI and Python assurance engine and a React, TypeScript, and Vite dashboard. The core verification path uses deterministic parsing and rules for questions that can be answered reproducibly, including citation resolution, metadata consistency, pinpoint validation, source hashes, and exact-text checks.
Official judgments are stored in immutable, versioned snapshots with numbered paragraph anchors and provenance. We introduced Case Maps to represent a judgment's structure, propositions, limitations, attribution, and review state. These fields are tiered by evidential reliability so that machine-proposed legal interpretations cannot silently trigger a hard gate or become benchmark truth.
For discovery, we built a separate offline candidate-authority path over a curated research catalogue. Its output is deliberately isolated from the verdict engine. For persistence, the project includes an optional Supabase schema with organisation boundaries, role-aware review controls, audit history, feedback, and versioned Case Map revisions; the demo can also run entirely from a frozen local snapshot.
We designed five source-locked demonstration scenarios covering a fake citation, mismatched case identity, an accurate quotation that does not support the claim, a non-holding presented as the court's view, and issue-specific negative treatment. This makes the system's reasoning visible in a short, repeatable demo.
Challenges we ran into
The biggest challenge was resisting false certainty. A failed lookup is not proof that a case was fabricated, a retrieved keyword match is not proof of legal support, and a distinguished case is not automatically bad law. We therefore had to design conservative verdict states and make abstention a useful product outcome.
Legal context also proved harder than document matching. Determining whether text is a holding, obiter, a submission, or minority reasoning can itself require legal judgment. We addressed this by attaching field-level provenance and human-review status to Case Map annotations, then restricting what unreviewed machine-generated fields are allowed to do.
Completeness created a different technical problem: there is no passage to cite for something missing from an answer. We designed a separate negative-findings format that discloses the corpus, issue tags, landmark set, retrieval configuration, and limitations behind each omission prompt.
Finally, we needed the demonstration to be fast and reliable without depending on live court websites, cloud models, or hosted infrastructure. Immutable snapshots, cached evidence, deterministic fallbacks, and an offline demo mode gave us a repeatable presentation path while preserving a route to a production architecture.
Accomplishments that we're proud of
- We delivered an end-to-end workflow from pasted AI answer to claim-level evidence, conservative findings, review flags, and a concise lawyer handoff.
- We modelled five distinct levels of legal hallucination instead of treating every problem as a fake-citation check.
- We made every important finding traceable to exact paragraphs, source status, hashes, versions, and decision rules.
- We separated benchmark gold, lawyer-approved annotations, machine-proposed labels, lexical ranking, and practitioner feedback so that weaker evidence cannot silently become legal truth.
- We created a five-scenario, source-locked demo and a fixed Singapore employment-restraint benchmark that can be run from a frozen local snapshot.
- We built human governance into the product: reviewer corrections create new versions, feedback enters a controlled review workflow, and uncertainty is escalated rather than hidden.
- We kept the product honest. ProofMark is a bounded evaluation tool, not legal advice, a comprehensive citator, or an autonomous legal-research system.
What we learned
We learned that trust in legal AI depends less on producing one impressive confidence score and more on exposing the chain of evidence behind each conclusion. Provenance, versioning, abstention, and a well-designed review path are product features, not administrative details.
We also learned that structured legal representations create both leverage and risk. A Case Map lets the system reuse expensive analysis across many audits, but an error in a cached legal interpretation can affect every later result. Field-level provenance and versioned human correction are therefore as important as retrieval accuracy.
Most importantly, we learned to match each question to the least probabilistic method capable of answering it. Deterministic checks should remain deterministic; semantic tools should operate inside evidence constraints; and genuinely judgmental questions should stay visible to lawyers.
What's next for ProofMark
Our next priority is to connect lawyer-approved Case Maps to every runtime audit so that role, modality, limitations, and factual fit are evaluated consistently. We will then strengthen the independent completeness path, expand the reviewed authority and treatment graph, and add statute version and currency checks.
Beyond the MVP, we plan to grow the expert-labelled adversarial benchmark, publish calibrated per-check performance, and add recomputation when an approved source or Case Map changes. Hosted Supabase deployment, durable background workers, licensed source integrations, and additional legal domains will follow only after their legal, security, evidence, and quality gates are met.
Built With
- ai
- ml
- natural-language-processing
- react
- supabase
- vite
Log in or sign up for Devpost to join the conversation.