Inspiration
Robot fleets record more video than anyone can watch, and the standard answer is "search it with text." We tested that assumption instead of accepting it, and it collapsed: cos("open the drawer", "close the drawer") = 0.9766. To a text encoder those are the same sentence. Across 57 kinds of moment, 24 typed queries landed at or below random chance and six returned nothing correct at all.
So the operator cannot describe the moment they want. They can only point at one. That reframes the whole problem: if a human has to look at clips anyway, the agent's job is not to find footage, it is to stop asking the human the same question twice.
What it does
Precedent is a video triage agent for robot fleets. It pulls unseen episodes off a queue, decides what each one is, and escalates to a human only when its own past decisions do not already answer the question.
Cold, it escalates 94% of its queue. With memory on, that falls to 38% while staying 96.2% correct on the 346 episodes it filed alone. With memory off it never improves: 100% escalation, forever. That gap is the entire submission.
Then the part a vector store cannot do. When an operator overturns one filing, the correction walks the precedent graph backwards and supersedes every decision that leaned on it: one overturn reached 16 direct dependants and 18 more transitively, 34 filings corrected in a single transaction.
How we built it
The memory is four tables in CockroachDB, not a bolt-on vector cache. inbox (task state, claimed atomically), filings (decisions and confidence), filing_precedents (the graph, which is what makes memory correctable), and verdicts (the only place a human writes).
Retrieval is two stages. Stage 1 is one SQL query over CockroachDB's VECTOR columns with distributed vector indexing, and it never scores 98.6% of the corpus. The indexed vector is concat(qf, qs)/√2 of the encoder's whitened pooled channels, so the database's cosine is the model's own prefilter rather than an approximation: verified to 2.3e-08. Stage 2 re-ranks the 48 survivors with DTW over full descriptor traces, which is the part a single vector cannot do because it sees the order things happened in.
The cascade is a recursive CTE walking the graph backwards, with tenancy enforced inside the recursion.
AWS: S3 holds the corpus (6,958 objects, 2.35 GB, private with SSE, GET-only CORS). Lambda runs the worker in us-east-2 beside the cluster because it makes only database round trips and reads zero S3; the deploy script parses that region out of the DSN rather than defaulting. EventBridge schedules it. Two IAM identities, because one lives in a secret store: the deployer can write, the app gets GetObject on two prefixes and nothing else.
The encoder (V-JEPA 2 plus SigLIP 2) is offline. Nothing in the request path does a forward pass over video.
Challenges we ran into
The first agent design made things worse. Storing memory increased escalation. The root cause was that it keyed on absolute similarity, which sat at 0.16 to 0.27 whether the answer was right or wrong. Ranking and calibration are different properties, and we had confused them. Rebuilding it on top-k consensus fixed it, and sweeping the threshold rather than picking one gave 0.5 → 8.0%, 0.6 → 81.8%, 0.8 → 96.2%.
Tenancy leaked through the cascade. A correction in one fleet was superseding another fleet's filings. Silent, and both a privacy and a correctness bug. A test caught it, not a review.
SKIP LOCKED is a trap on CockroachDB. Under SERIALIZABLE a freshly committed row can be briefly invisible, so a worker reads an empty claim as "queue empty" and stops with work pending. Nothing raises. Three failing tests found it.
An N+1 that only exists in production. Moving traces to S3 exposed that stage 2 fetched its 48 candidates one at a time inside the ranking loop: 49 sequential round trips, 37.3 s for a query whose real work is under a second. On local disk that cost is microseconds, so it had been invisible for the entire build.
The migration broke the thing it was migrating. Repointing the corpus to S3 rewrote a column the local deployment also reads, so running without AWS silently stopped playing video.
Accomplishments that we're proud of
The failure is measured, not asserted. Every result on screen is graded live from the query you just ran, against the folder the dataset filed the episode under, which retrieval never reads and never matches against.
We shipped a control arm. Memory-off is not a rhetorical device, it is a second run of the same agent over the same queue, and it flatlines.
The read path stopped holding the corpus. 17.64 GB resident became 1.37 GB, which is what let the whole thing run on a free CPU host, and we verified the compression was lossless where it counts: identical top-1 on 12 of 12 queries, identical ordering 12 of 12.
Eight tests pass against CockroachDB Cloud, not against Postgres pretending to be it.
What we learned
Which view to match on is a measured decision, and it flips. Appearance beats motion 0.786 to 0.391 on a corpus that varies in scene; motion beats appearance 0.748 to 0.649 on one that varies only in action. There is no default.
A database that expects retries is a contract, not a detail. 40001 handling belongs in one place so every caller gets it free. Omitting it is the single most common way a CockroachDB app is wrong in production, and it looks fine on a laptop.
Memory is only useful if it can be wrong. A memory you cannot correct is a liability that compounds. Storing the dependencies between decisions, not just the decisions, is what turns one human ruling into fleet policy and lets you take it back.
What's next for Precedent
Real footage. The corpus is 3,402 simulated kitchen episodes, and the retrieval claim is about text-versus-example on that corpus, not about generalising to your CCTV. We have 2,414 real robot clips already encoded and unused, which is the honest next test.
Multi-tenant beyond the test suite: fleets currently isolate correctly, but nothing enforces it at the connection level yet.
Stage 2 scales with trace length squared and the Sakoe-Chiba band has never been swept, so there is measurable latency sitting on the table.
And the operator loop should close: today a human overturns a filing, but the consensus threshold that decides when to escalate is fitted offline rather than moved by the corrections themselves.## Inspiration
Built With
- amazon-eventbridge
- amazon-web-services
- aws-iam
- aws-lambda
- boto3
- cockroachdb
- docker
- fastapi
- hugging-face-spaces
- huggingface
- javascript
- numpy
- pgvector
- postgresql
- prometheus
- psycopg2
- pytest
- python
- pytorch
- robocasa
- siglip-2
- sql
- transformers
- uvicorn
- v-jepa-2
Log in or sign up for Devpost to join the conversation.