Inspiration
Qualitative research — interview transcripts, survey free-text, field notes — is usually analyzed by hand or with expensive tools like NVivo. As a student doing real research (including prior work on WhatsApp misinformation studies at FUTA), I felt this pain directly: hours of manual highlighting and spreadsheet-building just to find themes that AI could surface in seconds.
What it does
Research Doc Intelligence takes raw research documents and:
- Extracts themes with supporting quotes from each document
- Synthesizes recurring themes and contradictions across multiple documents
- Exports a clean, citation-backed report as a downloadable Word document
In a test run on two real research documents about WhatsApp misinformation spread among university students, the pipeline correctly identified 4 recurring themes shared across both sources — trust in familiar senders, forwarding mechanics, protective sharing motivation, and a shared evidence gap — and explicitly confirmed no contradictions were found between them, rather than leaving that check ambiguous.
How I built it
- Backend: Django 6.1, Python 3.11+, with an ingestion pipeline that chunks documents (~1800-word chunks, split on paragraph/sentence boundaries) and sends them through a two-stage LLM pipeline
- AI:
zai-org/GLM-5.2via Featherless.ai (OpenAI-compatible API), orchestrated through a Hermes agent sandbox — one prompt for per-chunk extraction, a second for cross-document synthesis - Frontend: Django templates + Tailwind CSS (CDN, no build step)
- Export: python-docx handles both parsing and report generation — one library for both directions
- Caching: a filesystem cache (
responses/{sha256}.json) to avoid re-spending API credit on repeated dev runs
The pipeline has clean separation of concerns: ingestion → LLM client + prompts → orchestration → export, each in its own module.
Challenges I ran into
- Token limits truncating JSON mid-response. My initial budget (800/1200 tokens) was calibrated for "compact JSON," but GLM-5.2 produces detailed, high-quality output that needs more room — truncated JSON fails silently and looks like a successful call that returned nothing. Fixed by raising limits to 3000/4000 tokens.
- Bloated synthesis input causing empty responses. Sending full per-document extractions (confidence scores, flags, everything) into the synthesis step overloaded the model's effective attention on ~9KB of input. Trimming to just theme labels, descriptions, and quotes cut input size by a third and fixed it.
- Markdown fences in "JSON-only" responses. GLM-5.2 ignored the
"no markdown fences" instruction often enough that I had to build robust
fence-stripping plus a graceful fallback (a
parse_errorflag rather than a crash) so one bad chunk doesn't take down the whole pipeline. - Stale cached failures. An early empty API response got cached, so even after fixing the root cause, the pipeline kept returning the old empty result until I cleared the cache — a reminder that cache invalidation is exactly as hard as everyone says.
- A no-IPv6 network fighting the API client. My dev machine has no
IPv6 route, but
requests/urllib3tried IPv6 addresses first and failed instead of falling back to IPv4 like curl does — fixed by forcingAF_INETin the connection pool.
What I learned
Never trust an LLM to follow output-format instructions perfectly — always parse defensively and fail gracefully per-chunk rather than pipeline-wide. Also: input size to an LLM step isn't free just because it fits the context window — trimming to only what a step actually needs measurably improved output quality, not just cost.
What's next
- PDF support (text-based, not OCR)
- User accounts and saved analyses
- Interactive theme editing — let a researcher rename, merge, or reject LLM-suggested themes
- Streaming UI so results appear as each chunk completes, not in one batch
- Custom prompt templates so researchers can define their own coding scheme
Built With
- django
- featherless-ai
- glm-5.2
- hermes
- htmx
- llm
- natural-language-processing
- python
- python-docx
- qualitative-research
- sqlite
- tailwindcss
Log in or sign up for Devpost to join the conversation.