Inspiration
Scientific AI can produce polished answers while hiding a dangerous gap: a citation may exist without actually supporting the claim being made. Source verification, model interpretation, and human judgment are often collapsed into one opaque response.
EvidenceForge began with a question: what if every AI-assisted scientific decision had an inspectable chain from the exact source passage to the final experiment?
What it does
EvidenceForge is a human-governed scientific evidence workspace. A user first approves a bounded source packet and research scope. The system then:
- decomposes the question into testable claims;
- links every evidence card to an exact source passage;
- keeps deterministic verification, model-assisted assessment, and human review separate;
- identifies supporting, contradicting, and unresolved evidence;
- produces calibrated conclusions instead of invented confidence scores;
- selects the research gap that would most change the decision;
- proposes a bounded, falsifiable experiment;
- obtains critique from a second model family;
- applies only objections explicitly accepted by a human;
- preserves unresolved risks, failures, repairs, and the final human decision in a canonical JSON export.
Our demonstration asks whether a biodegradable battery could replace a coin cell in a single-use humidity sensor requiring 72 hours of operation. One source reports 77 hours—but only while the battery was unloaded. Another reports performance degradation after one hour due to drying. EvidenceForge does not average these into an overconfident “yes.” It identifies loaded 72-hour operation as the decision-changing evidence gap and proposes an experiment to test it.
How we built it
EvidenceForge is a Next.js, React, and TypeScript application with versioned Zod contracts.
The interface covers intake, evidence inspection, conclusions, gap selection, experiment planning, adversarial review, selective revision, final human approval, and export. Model providers sit behind server-side adapters, while deterministic application code owns provenance hashes, identifiers, validation, review state, and audit history.
The current live configuration uses Featherless with Mistral as the primary model and Qwen as the reviewer. The reliable hackathon demonstration uses deterministic fixture playback so judges can inspect the complete workflow without depending on provider latency or availability.
Challenges we faced
The hardest challenge was keeping model-generated content from controlling application-owned facts. Models should propose scientific interpretations—not fabricate provenance, verification results, review decisions, or execution metadata.
We introduced compact model-facing schemas, strict local validation, deterministic hydration of application-owned fields, bounded repair attempts, and append-only failure records.
A bounded live run successfully completed evidence extraction, entailment, and synthesis, but experiment planning failed schema validation. We preserved that failure instead of presenting it as a successful end-to-end run, and retained fixture mode as the honest, deterministic demonstration.
Accomplishments
- Built the complete claim-to-experiment workflow.
- Preserved exact source passages and separate verification layers.
- Implemented selective, field-level revision from human-accepted objections.
- Produced byte-stable canonical exports.
- Preserved failures and repairs instead of silently overwriting them.
- Passed 642 unit tests, 291 evaluation tests, and 64 Chromium browser and accessibility journeys.
- Published a privacy-reviewed Software Development release to GitHub.
What we learned
Structured model output is not enough to make scientific AI trustworthy. The application must own provenance, invariants, audit history, and human authority.
The most valuable result is not a more confident answer. It is a decision record that shows what the evidence supports, what remains unresolved, and which experiment could change the conclusion.
What's next
Next steps include improving live experiment-planning conformance, testing additional scientific domains, and running a properly controlled comparison against a strong single-prompt baseline. We will continue keeping fixture evidence, live evidence, and measured evaluation clearly separated.
Built With
- ai
- api
- eslint
- featherless
- github
- json
- mermaid
- mistral
- next.js
- node.js
- playwright
- pnpm
- qwen
- react
- rest
- typescript
- vitest
- zod

Log in or sign up for Devpost to join the conversation.