Inspiration
Inspiration
We kept seeing the same pattern: researchers trust a claim because a paper cites something, or because a chatbot summarizes something—without checking whether the evidence actually supports it.
Literature tools are great at search and summary. They are weak at verification. They lean toward supporting quotes, rarely hunt for contradictions, and usually return one opaque answer with no audit trail. Writing happens in an editor; checking happens somewhere else. Context breaks.
We wanted a tool that asks a harder question:
Not “What do the papers say?”
But “What does the evidence prove?”
That idea became Papel AI.
What it does
Papel AI is a Word-like research manuscript editor with multi-agent claim verification built in.
- Detect — As you write or load a template, it finds scientific claims (comparative, quantitative, causal, limitations), scores them, and logs them with highlights.
- Verify — On Verify, agents decompose the claim, run support + adversarial queries, retrieve from Semantic Scholar and arXiv, extract evidence, critique conflicts, and return a verdict: Supported, Partial, Contradicted, or Insufficient.
- Explain — Results show scope, evidence strength, why a claim is not fully supported, optional safer wording, and a full Show Me Why trace.
How we built it
We built a React + Vite app with a TipTap/ProseMirror manuscript surface (multi-page sheets, templates, claim log, agent panel).
- Phase 2: staged claim pipeline — clean → segment → candidate filter → LLM classify → calibrate/dedup
- Phase 3: multi-agent verification — decompose → query expand → retrieve → rank → evidence + adversarial critic → verdict engine
- Phase 4: explanation UI — scope, strength, alternatives, Show Me Why
Stack: TipTap, Groq/Gemini for classification, Semantic Scholar + arXiv for retrieval, hybrid ranking, client-side PDF text path where needed.
Challenges
- Claim detection quality — scientific text is messy (abbreviations, decimals, citations). We had to tune filters so we catch real claims without drowning in noise.
- Support bias — a single “helpful” LLM call always leans positive. We forced an explicit adversarial channel so contradictions are first-class.
- APIs in a live demo — rate limits (e.g. 429), CORS, and abstract-only papers. We treated Insufficient and clear scope as valid outcomes, not failures to hide.
- Multi-page manuscripts — UI pages vs full-text analysis. We learned to separate visual pages from analysis text so detection can see the whole paper without breaking the editor animation.
What we learned
- Verification UX matters as much as model quality. Users need scope and “why,” not only a green badge.
- Multi-agent helps when roles are real: find, extract, attack, decide—not five agents saying the same thing.
- Honest systems must be allowed to say “not enough evidence.”
- Shipping a demo teaches more than architecture slides: prompts, rate limits, and edge cases show up immediately.
What’s next
Deeper full-text evidence, stronger page provenance, and verification run history so the same claim can be re-checked and compared over time.
Papel AI — detect the claim, stress-test it, explain the outcome.
Log in or sign up for Devpost to join the conversation.