Inspiration
Two problems we kept running into at Hopkins, from opposite ends of the same pipeline.
On the research side, a group's knowledge is scattered across papers, protocol drafts, analysis memos and the heads of a few people. Evidence gaps and contradictions go unnoticed, and questions go to whoever is nearest rather than whoever knows.
On the trial side, eligibility criteria are written blind. Nobody knows how many real patients survive an inclusion/exclusion list until the trial is open, and by then it is too late: roughly 80% of trials miss their enrollment timeline and about a fifth of sites enroll no one. Screening is still a one-time chart review, so a patient whose eGFR crosses a threshold next Tuesday is simply never seen.
Cohort connects the two: find what the evidence is missing, design the trial, and let the trial find its own patients.
What it does
Cohort is one graph running in both directions.
Research. Papers (PubMed / OpenAlex), protocols, internal documents and people are ingested into a knowledge graph. Deterministic detectors find evidence gaps (an intervention studied in CKD stage 3 but never stage 4), contradictory claims, stale guidelines, zombie trials with no published results, and bus-factor risks where one person holds all the edges. "Ask" answers questions with citations, and when the knowledge base cannot, it routes the question to the right expert and writes their answer back as a reviewed claim.
Design. From a gap, an agent drafts a trial concept (PICO plus criteria). Every criterion is tested live against the patient database: an attrition waterfall shows which criterion kills enrollment, sensitivity sliders recompute the cohort in under 300 ms, and a representativeness panel compares the surviving cohort to the population.
Recruit. Published criteria become standing queries. Free-text eligibility is parsed into a predicate DSL and executed as SQL over a TimescaleDB hypertable of observations. When a new lab arrives, only the affected patient × trial pairs are re-evaluated. A patient flipping eligible pulses on the graph, the coordinator gets a spoken brief, a multilingual voice agent pre-screens the patient, the conversation becomes a proposal, and a human approves enrollment. The funnel updates from continuous aggregates.
How we built it
- Data: 10,000 Synthea patients generated for Baltimore (2,000 on the demo database), 100 MIMIC-IV demo patients and 6 published case-report vignettes for validation, 51 curated recruiting trials pulled live from ClinicalTrials.gov, 754 real PubMed papers linked by NCT id.
- Database: Tiger Data (TimescaleDB Cloud) as the single store: hypertables for observations, eligibility events and funnel events; continuous aggregates for the live charts; compression; pgvector for 768-d embeddings of patients, trials, papers and document chunks.
- Matching engine, three layers. (1) Structured predicates compiled to SQL, deterministic and incremental. (2) A semantic judge (Gemini over retrieved patient notes) for criteria that resist structure, which can only resolve unknown and can never override a hard fail. (3) An embedding prefilter and ranking. Confidence is calibrated with Jev; on our labeled pairs the expected calibration error is $\text{ECE} \approx 0.06$. Eligibility is fail-closed:
$$ \text{status} = \begin{cases} \text{excluded} & \exists\, c \in E:\ c = \text{fail} \ \text{eligible} & \forall\, c \in I:\ c = \text{pass} \ \wedge\ \neg\exists\, c \in E:\ c = \text{fail} \ \text{unknown} & \text{otherwise} \end{cases} $$
- Voice: ElevenLabs TTS for coordinator briefs; an ElevenLabs Conversational Agent with client tools so the pre-screen conversation writes structured answers, not just audio.
- Memory and reports: Backboard threads per patient and coordinator with RAG over protocol documents; Snowflake Cortex for site-level reports.
- App: React 19 + Vite, a d3-force canvas graph ported from Corpus, recharts, Express 5 with socket.io for live events, a replay engine that streams held-out observations at an accelerated clock. Dockerized for DigitalOcean; marketing site on Vercel.
- Evaluation: a golden suite for criteria parsing and eligibility precision/recall runs with
npm run eval; 459 unit tests across server and client.
Challenges
- Synthetic time. Synthea's data ends in 2020, so every "within 365 days" lab window read as unknown. We time-shift the entire dataset at load and hold back the last 180 days for replay.
- Compressed chunks are slow for "latest value per patient". Our eligibility queries collapsed under compression until we added a trigger-maintained latest-observation table. Full re-evaluation of 10k patients now takes about 90 s; incremental updates are milliseconds.
- Free-tier limits everywhere. Gemini and Jev credits ran out mid-build and Tiger's 750 MB cap forced downsampling. Every provider now has a real fallback (rule-based parser, hashed embeddings, scripted pre-screen) and hourly spend guards, so the demo never depends on a key.
- Keeping the LLM honest. The most important design decision was that no model output alone can mark a patient eligible or trigger contact. The deterministic layer decides; the model only explains and resolves unknowns; a human signs every proposal.
- Scope. We rebuilt the UI once after judging the first graph-centric version too thin for real decisions; the final workspace is organized around what a coordinator actually does: protocol, feasibility, screening worksheet, patient 360.
What we learned
Eligibility criteria are a query language nobody wrote a compiler for. Once they are predicates, feasibility, recruitment and monitoring are the same computation run at different times. Time-series Postgres made that trivial to keep live. And the research half is not decoration: the trials worth running are the ones the evidence graph says are missing.
What's next
FHIR ingestion from a real EHR sandbox, IRB-facing audit exports, and site-to-site federation so a rare-disease trial can find its twelve patients across hospitals without moving data.
Built With
- backboard
- d3.js
- digitalocean
- docker
- elevenlabs
- express.js
- framer-motion
- google-gemini
- graphology
- node.js
- openalex
- pgvector
- postgresql
- pubmed
- react
- recharts
- socket.io
- synthea
- tiger-data
- timescaledb
- typesafe
- typescript
- vite
- zustand
Log in or sign up for Devpost to join the conversation.