The Problem: Academic Research is Broken
Writing a literature review today requires researchers to download dozens of PDFs, manually copy-paste abstracts into a chatbox, and beg the model to summarize them. The result? The LLM hallucinates methodologies, loses track of citations, and breaks LaTeX math formatting.
I realized that LLMs don't need better prompts—they need ==Agent-Native infrastructure==.
Why this is a perfect fit for WebMCP
Before WebMCP, LLMs were blind to the user's UI. By implementing WebMCP, I exposed 13 strictly typed tools directly to the agent. This creates a massive UX paradigm shift:
- [x] Zero Copy-Pasting: The agent fetches full-text HTML directly via
extract_findings. - [x] No More Hallucinations: The
compare_paperstool forces the agent to read structured matrix data. - [x] Synchronous Collaboration: When the agent calls
draft_section, the manuscript in my React UI updates instantaneously.
How I Built It: The Core Mechanisms
I built PaperPilot using Next.js 14, React 18, and TypeScript, deployed seamlessly on Vercel. To make this production-ready, I engineered three core mechanisms:
1. The Zero-Backend Event Bus
To achieve real-time UI synchronization without the latency of WebSockets, I implemented a decoupled event architecture. When ChatGPT executes a mutation tool:
document.modelContext.registerTool({ name: "draft_section" })
It triggers a pure async Next.js Server Action. This action mutates the local data store and emits a native DOM CustomEvent (e.g., paperpilot:outlines-changed). The React UI listens to this bus and re-renders instantaneously.
2. Heuristic Extraction & Grammar-Aware Parsing
Academic HTML is notoriously messy. To prevent context window pollution, I engineered an extraction algorithm that isolates methodology while discarding boilerplate (Appendices, References). The extraction logic \(\mathcal{E}\) for a document \(D\) composed of sentences \(s_i\) is defined as:
$$ \mathcal{E}(D) = \bigcup_{i=1}^{N} s_i \quad \text{subject to} \quad \text{class}(s_i) \neq \text{"ltx_bibliography"} $$
Furthermore, naive sentence splitting ruins context by breaking on abbreviations like "e.g." or "et al." I parse sentences using a strict, grammar-aware Regex formula to guarantee abbreviation bypass:
(?<=[a-z]{3,}[.!?])\s+(?=[A-Z0-9])
3. Math-Safe LaTeX Compilation
ChatGPT frequently generates inline mathematical formulas, such as \(\mathcal{O}(\log d)\). If I blindly escaped reserved LaTeX characters (like _ or %), it would break the math mode syntax. I engineered a custom AST-like parser that isolates \($\) and \($$\) blocks, safely escapes rogue characters only in standard text strings, and then restores the math blocks.
The agent also autonomously injects \cite{} markers linked to a dynamically generated \bibitem matrix, guaranteeing ==100% compilable Overleaf code==.
The Proof: My 20-Paper Stress Test
To prove PaperPilot isn't just a brittle hackathon demo, I built an automated stress test against 20 seminal AI papers (including Attention Is All You Need and Mistral 7B).
| Metric | Success Rate | Truncation Errors |
|---|---|---|
| Methodology Extracted | 20 / 20 | 0 |
| Key Claims Extracted | 20 / 20 | 0 |
| Limitations Extracted | 20 / 20 | 0 |
| Conclusion Extracted | 20 / 20 | 0 |
What I learned & What's next
I learned that strict typing in WebMCP inputSchemas is the absolute key to preventing agent hallucinations. Moving forward, I plan to expand my WebMCP toolset to integrate directly with Zotero and Mendeley, allowing researchers to bring their existing private libraries into the Agent-Native web.
Built With
- agents
- api
- arxiv
- automation
- chatgpt
- front-end
- git
- github
- javascript
- latex
- llm
- localstorage
- lucide-icons
- next.js
- openai
- parsing
- react
- regex
- research
- server-actions
- tailwind-css
- typescript
- vercel
- webmcp
- zero-backend
Log in or sign up for Devpost to join the conversation.