We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Having previously explored and built complex cognitive memory architectures for autonomous agents, we realized a critical bottleneck in current AI tools: memory and context are only as good as the facts fed into them. While tools like Perplexity Pro and OpenAI Deep Research are incredibly powerful, there is a massive gap for a fully open-source, production-grade alternative that developers can host locally. We wanted to build a system that doesn't just return a quick, single-shot LLM answer, but actually "thinks" like a human researcher—recursively planning, digging through primary sources, verifying claims, and catching contradictions, all while remaining resilient enough to run completely offline if needed.

What it does

ResearchForge AI is an autonomous, multi-agent knowledge engine that transforms a high-level query into a publication-ready intelligence report. When given a complex prompt, the orchestrator breaks it down into targeted hypotheses and search vectors. It conducts live web scraping across multiple sources, extracts high-signal text, and embeds it into a hybrid vector/lexical memory system. The core innovation is our rigorous fact-verification engine, which cross-examines multiple independent domains to detect contradictory evidence and assigns confidence scores. The entire thought process, including active web queries and discovered sources, is streamed to the user in real-time via Server-Sent Events (SSE) on a clean, Perplexity-inspired dark UI.

How we built it

We architected the platform with a clear separation of concerns to ensure scalability and local resilience:

Frontend: Built with Next.js 14 (App Router), TypeScript, Tailwind CSS, and Framer Motion for a distraction-free, highly responsive interface.

Backend & Orchestration: Powered by Python 3.11+ and FastAPI, utilizing AsyncIO for high-performance concurrent web retrieval and SSE streaming.

AI Engine: The primary reasoning layer utilizes Google Gemini 1.5 Flash, backed by a seamless fallback to local Ollama inference (qwen3:1.7b) and a deterministic heuristic synthesizer for zero-setup offline capability.

Web Retrieval & Memory: We integrated Tavily AI and DuckDuckGo for search, paired with a Playwright headless browser for deep DOM extraction. The retrieved data is synthesized into a hybrid memory index using FAISS (for semantic vector search) and BM25 (for keyword matching), managed via SQLAlchemy.

Challenges we ran into

One of the toughest hurdles was orchestrating the recursive inquiry loops without getting trapped in infinite search spirals or hitting rate limits. Managing the headless Playwright browser instances concurrently required strict timeout protocols and automatic cache cleanup to prevent memory leaks during deep web extraction. Additionally, designing the real-time SSE telemetry to stream agent "thoughts" and citations reliably to the Next.js frontend, while simultaneous background tasks were scraping and embedding text, required meticulous async state management.

Accomplishments that we're proud of

We are particularly proud of the hybrid fallback architecture. Research tasks never outright fail; if external API quotas are exhausted or the system is air-gapped, it seamlessly degrades to local Ollama inference without disrupting the user experience. We also successfully built an automated cross-examination engine that genuinely flags conflicting claims across different web sources rather than hallucinating a false consensus.

What we learned

We gained deep, practical insights into managing real-time data streams (SSE) across a decoupled stack. Fine-tuning the hybrid retrieval pipeline taught us how powerful combining traditional lexical search (BM25) with vector embeddings (FAISS) can be for grounding LLMs in reality. We also learned how to better structure multi-agent recursive workflows to optimize for both speed and rigorous accuracy.

What's next for ResearchForge AI

We plan to expand our local offline capabilities by introducing support for larger models like Llama 3 and Mistral out-of-the-box. Next on the roadmap is integrating PostgreSQL (via pgvector) for enterprise-scale memory persistence, replacing the local SQLite/FAISS setup for heavier workloads. Ultimately, we want to package the web retrieval engine into a standalone browser extension, allowing ResearchForge to act as a proactive, on-page fact-checker for users as they browse the web

Built With

Share this project:

Updates

Submission history