Inspiration
Research is slow because the hard part isn't writing — it's the grind before writing: finding what's already known, spotting where it's thin or contradictory, and designing something worth testing. We wanted to see whether an agent could do that grind itself, not just summarize a paper someone hands it.
What it does
Give VERITAS a research question. It autonomously discovers literature (Semantic Scholar + Crossref), synthesizes a state-of-the-art summary, identifies candidate research gaps and contradictions, and proposes future directions — all grounded in what it actually retrieved, not invented. If a dataset is attached, it goes further: it designs and runs real statistical/ML experiments with scikit-learn and SciPy, critiques its own evidence, and decides whether to run another iteration before finalizing. It ends by producing a full research proposal — methodology, experimental plan, evaluation metrics, reproducibility plan, and references — downloadable as Markdown or PDF.
How we built it
Gemini 3.5 Flash reasoning is orchestrated through Google ADK's LlmAgent + Runner, with specialist sub-agents for literature synthesis, methodology, and critique. A persistent ResearchState (SQLite locally, Firestore optionally in the cloud) tracks the investigation across its full lifecycle. The orchestrator runs a real decide → experiment → critique loop, not a single LLM call — deterministic Python (scikit-learn, SciPy) handles every numeric claim, so evidence is computed, not hallucinated. The frontend is Streamlit; a FastAPI backend exposes the same pipeline programmatically. Cloud Run, Pub/Sub, and Cloud Storage support optional cloud-mode deployment, with a clean local fallback for development.
Challenges we ran into
Free-tier Gemini rate limits (as low as 5 requests/minute on some models) meant the multi-agent pipeline could exhaust quota mid-run; we mitigated this by moving to a Flash-Lite model and keeping iteration counts low for demos. We also hit a subtler reliability issue: background research jobs ran on daemon threads tied to the local process, so anything that killed the process (a restart, a sleeping machine) silently orphaned the job with no error — we added a watchdog that auto-fails any job with no progress for over two minutes, so a stall is now immediately visible instead of appearing to hang indefinitely.
Accomplishments that we're proud of
The autonomous loop is real, executable control flow — literature retrieval, gap analysis, experiment design, deterministic computation, self-critique, and a genuine decide-to-iterate step — not a single prompt dressed up as a workflow. Numeric findings are never invented by the LLM; they come from actual scikit-learn/SciPy computation, which we treat as a hard line the system won't cross.
What we learned
Draft below, grounded in what actually happened while building this — edit freely, this should sound like you:
We learned that most of the hard engineering in an "autonomous agent" isn't the LLM call — it's everything around it. State persistence across background threads, handling rate limits gracefully instead of failing silently, and making sure a stalled job is visible rather than invisible all mattered more to reliability than the quality of any single prompt. Debugging a daemon thread that dies silently when a process restarts taught us to treat "is it actually running" as a first-class question the UI has to answer honestly, not just assume.
What's next for VERITAS — Autonomous Research Scientist
Tighter dataset-to-topic matching so the experiment stage engages more often instead of gracefully skipping; broader literature source coverage beyond Semantic Scholar/Crossref; and hardening the Cloud Run + Pub/Sub path for genuinely long-running, multi-day research jobs rather than single-session runs.
Built With
- docker
- fastapi
- firestore
- gemini
- google-adk
- google-cloud
- google-cloud-pubsub
- google-cloud-run
- pandas
- python
- scikit-learn
- scipy
- sqlite
- streamlit

Log in or sign up for Devpost to join the conversation.