-
-
Title — Unattended Taskmaster agent. Living evidence graph for LLM RAG. Demo: Keytruda + NSCLC.
-
Problem — LLMs invent citations and trial IDs that look real.
-
Solution — Search public sources, write a checkable graph, ground the LLM on those edges.
-
Live — Cloud Run demo. GET /health returns 200.
-
Graph — Frozen public demo: 14 nodes, 10 edges, 7 source families.
-
Compare — KEYNOTE-888 trap: Bare invents an ID; grounded and strict stay on the graph.
-
Digest — Unattended change digest: what, why, and ids.
-
Push — One-click push binds the graph into the LLM session (grounded or strict).
-
Safety — No causation, not clinical advice, public data only, no PHI.
Inspiration
LLMs invent trial IDs, paper citations, and over-claim numbers. In literature-heavy work that looks like evidence. We wanted an unattended agent that keeps a living, trust-scored evidence graph and grounds answers on edges you can open.
What it does
Living Evidence Graph is a Taskmaster: you set a research goal, then it runs by itself.
Demo goal: pembrolizumab / Keytruda + NSCLC, public APIs only. It fetches Open Targets, ChEMBL, ClinicalTrials.gov, PubMed, DailyMed, Europe PMC, and openFDA FAERS (reports, not rates), builds a drug–target–disease spine plus corroboration edges, scores trust, and emits a change digest (what / why / ids).
That graph is the RAG layer for Gemini. Bare can invent a pairing (demo trap: KEYNOTE-888 is not in the 14-node / 10-edge graph). Grounded cites openable edges. Strict answers graph-backed clauses and leaves gaps empty.
The live Cloud Run site is the public Keytruda / NSCLC graph only. Personal and enterprise use the same engine on a local folder (README: Clone and run → Personal / Enterprise). Those graphs stay on the operator’s machine and are not hosted on the public demo.
Live: https://living-evidence-graph-892760629727.us-central1.run.app Repo: https://github.com/qxiong888/living-evidence-graph Video: https://youtu.be/P9jyki4d4P8
How we built it
Google ADK + Gemini 3.5 Flash (Gemini API / AI Studio, not Vertex as the default). FastAPI on Cloud Run (us-central1, min instances 0). Cloud Scheduler to POST /scheduler for daily refresh (paused so the demo stays frozen). Optional Firestore. Public fetchers only: no scraping, no invented NCT/PMID/setid/ChEMBL/Ensembl IDs.
Challenges we ran into
Keeping the live demo identical to the video (baked 14/10 graph; no mid-judge refresh). Making grounded vs strict actually refuse missing trials instead of sounding fluent. Labeling FAERS as reports, not rates. Saying clearly that the graph is a retrieval layer, not fine-tuning, and not clinical advice.
Accomplishments that we're proud of
A closed unattended loop: goal in, daily public fetch, trust-scored living graph, change digest, then grounded or strict RAG with one-click push into the LLM. Live Cloud Run proof with a .run.app URL. A mixed question where KEYNOTE-888 is absent: bare invents a pairing, grounded cites edges, strict leaves that clause empty.
What we learned
A small, dated, auditable graph beats a larger fluent guess. Unattended refresh only helps if the digest and RAG stay honest when the graph is empty or a claim is missing.
What's next for Living Evidence Graph
Personal and enterprise private libraries (already in the repo) after the contest. No diagnosis product. The public Keytruda graph stays frozen until results.
Not a medical product. Not clinical advice. openFDA FAERS = voluntary reports, not incidence rates, not causation. Not endorsed by FDA, NLM, NIH, NCBI, Open Targets, or ChEMBL.

Log in or sign up for Devpost to join the conversation.