Inspiration

Before this hackathon, we spent a long design effort — 114 recorded architecture decisions — on an evidence-and-compliance layer for AI in regulated industries, built on one thesis: a system should never assert what it cannot prove. When TechJam's Track 1 asked for the middleware that the Agent Launchpad is missing, we recognized the same hole at hackathon scale: an agent will act on whatever reaches its context. Nothing stops restricted data from leaking into a prompt, nothing distinguishes a model's confident hallucination from a sourced fact, and when a trusted input later turns out to be wrong, nobody can say which decisions were built on it. So we distilled that design thinking into Walnut Rewind.

What it does

Walnut Rewind sits under the Launchpad as four small TypeScript modules:

  • No fact without a source. An agent can only propose evidence — a claim plus the exact place in a source file it came from. The middleware verifies every quote by exact byte match (source[start:end] === quote) and rejects on mismatch regardless of model confidence. Rejected text never enters the record — only the proof of its rejection.
  • Authorization before context. Permissions are checked before the prompt is built; each run executes against a sealed, hashed Context Capsule. In the demo, a planted tracer in the denied payroll claim appears zero times in anything the model saw — proven, not promised.
  • A tamper-evident flight recorder. Every step lands in a per-run, redact-before-persist, hash-chained ledger. Corrupt a copy at any entry and verification names the exact broken one.
  • Rewind. Mark one source compromised and the dependency graph flags exactly the runs and artifacts built on it — nothing unrelated. Recovery mints a new run with the bad fact denied; the old run is marked RECOVERED with its capsule hash and chain byte-identical. History is never rewritten.

The 3-minute demo (https://youtu.be/PGvlZ4ji-No) plays this as one incident: a hallucinated claim rejected, a payroll leak denied, a trusted launch date compromised, traced, and recovered.

How we built it

Under a strict written rulebook: frozen interface contracts before code, invariant tests before features, adversarial review at every gate, and a non-negotiable truthfulness rule — no artifact may state a result that isn't reproducible from a tracked file in the repo. The process practices what the product preaches: every claim in this description traces to a committed evidence artifact under results/.

Challenges we ran into

  • The free model quota died mid-recording (HTTP 429 at 11:53 PM, deadline at noon) — we switched endpoints and re-drove the demo state through the live pipeline.
  • Append-only bites back: you cannot "reset" a recovered run for a retake — by design. Every re-take meant rebuilding the incident from scratch through real runs.
  • Honesty is harder than features: we cut an entire feature (evidence-pack export) rather than stub it, and the README's limitations section states plainly what the runtime stream cannot observe.
  • Clarity, the hard way: live rehearsal kept exposing gaps between what the UI showed and what a newcomer could understand — the narration, one UX fix (reconcile now auto-refreshes the chat), and two rejected-claim explanations all came out of a human repeatedly asking "how would a stranger know that?"

Accomplishments we're proud of

npm run check exit 0 — 233 tests in 27 files; 0 npm audit --omit=dev vulnerabilities; a 12-stage end-to-end test of the whole thesis run 3× with no flake; a real-Chrome UI pass (24 screenshots, 0 console errors); blocking a run on conflicting evidence costs 0 model tokens (typed clarification instead of a silent guess); and a git history verified clean of secrets end to end.

What we learned

Enforcement only works before context construction — after the model has seen the text, every control is advisory. Byte-exact verification beats confidence scores. Append-only recovery is both simpler and more auditable than mutation. And a demo script improves fastest when a non-expert keeps asking "what does this mean?"

What's next

Real delegated tokens, ReBAC authorization, connector-based provenance, signed evidence-pack export, durable workflows — extension points are documented in README §16; none are dependencies of this POC.

Try it (local, one command)

npm install
ARK_API_KEY=<key> ARK_MODEL=ep-<id> npm run poc   # → http://localhost:3000
npm run check

Repository: https://github.com/MEHUL-MODI-Git/Nexus_TechJam_2026_Track_1_Walnut_Rewind

Where each Track-1 deliverable lives

  • 3-minute demo — https://youtu.be/PGvlZ4ji-No (real Agent Runs; failure, denial, and recovery cases live)
  • One-page architecture diagram (middleware, data flow, trust boundary, enforcement/instrumentation/recovery points) — README §4.1
  • Setup — README §10 · Problem & rationale — README §§1–3 · Design summary — README §4 · Automated tests — README §12 · Demo steps — demo/FULL-WALKTHROUGH.md · Limitations — README §15 · No secrets — README §17 (incl. the verified full-history scrub)
  • §1.10 optional-evidence checkboxes, all four — README §18, each mapped to where it is demonstrated and tested

Built With

Share this project:

Updates

Submission history