Inspiration

Professional research often stalls between an AI-generated answer and a decision-ready handoff. A colleague needs exact source context, unresolved questions, and a record of what was actually reviewed.

What it does

EvidenceDesk takes a focused research brief and public-source excerpts, invokes a local Strands agent to search evidence and record findings, validates quotation anchors against sealed snapshot paragraphs, and produces a portable review packet. Every input question remains present, including failures and gaps. Reviews are attributed to their human or automated author and appended separately from original model output.

How we built it

Python 3.12, Strands Agents SDK 1.54.0, a real Ollama/Qwen3 local model, FastAPI, an accessible browser interface, deterministic citation checks, atomic local persistence, and portable ZIP exports. Strands owns the actual model/tool loop. Each question has its own conversation and bounded model-call budget. No cloud inference or paid provider fallback is configured.

Challenges and lessons

A small model sometimes prints an answer without calling the recording tool. EvidenceDesk retries once, then validates strict JSON through the same checker, recording that fallback distinctly from a model-issued tool call. Questions without a valid finding remain unresolved. Exact quotation matching checks text, not the correctness of its interpretation.

Demonstrated results

The genuine local run used two attributed excerpts from official Strands documentation. Q1 and Q2 were recorded through model-issued tool calls; Q3 remained unresolved and used the separately recorded JSON fallback. The run made nine model calls in 24.68 seconds; that is one observation, not a performance guarantee. The review note is attributed to Codex automated review with decision needs_review. No human approval is claimed.

Thirty local tests passed. Checks verified all eight entries in the actual browser-downloaded packet manifest, the original-run consistency seal, and all three quotation anchors. Hashes establish byte consistency, not publisher identity or correct interpretation. The video uses genuine screenshots from this run and synthetic narration; it is not continuous recording or live inference.

Try it

MIT source: https://github.com/ZackaryLoevseth/evidencedesk-strands

Downloadable test build: https://github.com/ZackaryLoevseth/evidencedesk-strands/releases/tag/v0.1.0

Follow the README Python/Ollama setup, start evidencedesk serve, open the local page, load the real-source example, and run it. The local model is a separate free download. This is a local single-user application; no hosted workspace is supplied.

Next steps

Richer paragraph retrieval, source-version comparison, independent reviewers, and an authenticated hosted workspace are future work.

AI and third-party disclosure

OpenAI Codex assisted implementation, tests, documentation, and demonstration preparation. Application inference used the real Strands/Ollama stack. No prior application implementation was imported. Dependencies retain their licenses and source passages are attributed. Original application code is MIT licensed. No measured productivity improvement, guaranteed source accuracy, accepted entry, or award is claimed.

Built With

Share this project:

Updates

Submission history