Inspiration

Scientific papers contain results that shape real decisions, but reproducing even one figure can take days. The necessary information is scattered across prose, equations, captions, code repositories, and unstated conventions. General-purpose AI can make this worse by producing code that looks convincing without proving that it matches the paper. I work in theoretical quantum physics, where checking assumptions, parameters, and numerical behavior is everyday work. I built Q-Replicate to answer a smaller and more honest question: what is the first useful part of this paper that an agent can recreate with evidence, and what still requires a human expert?

What it does

Q-Replicate turns one measurable result from an arXiv paper into a checked, rerunnable computer experiment. The user pastes an arXiv link. A Strands agent reads the public paper record, matches the paper to an independently validated model pack, and chooses a bounded reproduction path. Deterministic tools then run the model and check that the result completed, stayed within sensible physical bounds, and agreed with a trusted reference. The user receives a chart, Jupyter notebook, evidence report, and an explicit next action for expert review. If no validated model pack matches, Q-Replicate stops safely. It still explains what it understood, but it does not invent equations, parameters, or a numerical result. The main demo uses the familiar example of hot coffee cooling toward room temperature. A second demo shows energy transfer between two quantum batteries. This makes the safety pattern understandable to a general audience while showing that it also applies to specialist research.

How we built it

The control plane is a Python agent built with the Strands Agents SDK. Its system prompt enforces an evidence-first sequence and gives it four narrow tools:

  1. read_arxiv_metadata
  2. select_validated_model
  3. inspect_validation_record
  4. build_reproduction_plan The service uses the Amazon Bedrock AgentCore application interface and can be deployed to AgentCore Runtime. The responsive web application is built with React, TypeScript, and Vinext. It retrieves public records directly from arXiv, uses GitHub's public repository search for clearly labeled code leads, runs deterministic simulations, and exports SVG, Markdown, and Jupyter artifacts. The coffee model is checked against an independent RK4 integration. Three quantum model packs are checked against QuTiP 5.3.1. The web layer never treats agent prose as numerical evidence; it renders only structured tool output.

Challenges we ran into

The hardest design decision was deciding what not to automate. “Reproduce any paper” is an attractive demo promise but a scientifically unsafe one. I narrowed the task to validated model families and made the unsupported path a first-class product outcome. Another challenge was communicating scientific provenance to non-specialists. The interface separates paper-sourced values from benchmark values, puts the plain-language result first, and keeps assumptions, checks, and limits one click away.

Accomplishments that we're proud of

  • A real safe-stop path: unsupported papers do not receive fabricated code.
  • Independent numerical validation for every executable model pack.
  • Downloadable evidence artifacts instead of an answer that disappears in chat.
  • A general-audience coffee demo and a specialist quantum-physics demo using the same agent policy.
  • A clean separation between probabilistic agent judgment and deterministic scientific checks.

What we learned

Agent reliability comes less from a longer prompt than from better boundaries. Narrow tools, structured outputs, independent references, and visible human-review gates make an agent more useful and easier to trust.

What's next for Q-Replicate

Next I would add expert-authored reproduction packs for more fields, ingest full-text equations and figure captions, connect paper-specific parameter mapping to a review queue, and evaluate the model-selection and safe-stop behavior on a larger benchmark set.

Built With

Share this project:

Updates

Submission history