Inspiration
Everyone has received a piece of mail they couldn't fully understand: a medical bill full of CPT codes, a parking ticket with escalating fines buried in legalese, a lease notice hiding an auto-renewal trap. The real cost isn't confusion — it's missed deadlines, overpaid fees, and unchallenged denials. The burden falls hardest on students, immigrants, and anyone reading in a second language.
Generic chatbots can define jargon, but they can't read your document, find your deadline, or draft your reply. We wanted to close that last mile — from "what does this say?" to "it's handled."
What it does
Unfog is a paperwork decoder. Snap a photo (or drop a PDF, or paste text) and get:
- 📖 A plain-language summary + one-line ELI5 of what the document actually means
- 🔑 Key facts extracted — and visually grounded: hover any fact and it highlights where that text sits on your document
- ⏰ Deadlines with countdowns and one-tap calendar (.ics) export
- 🛡️ A legitimacy check — it flags scam patterns (gift-card demands, arrest threats, fake URLs) before you pay a fake ticket
- ✅ An actionable checklist where every step has a "how do I do this?" button
- ✍️ Ready-to-send drafts — dispute letters, phone scripts, questions to ask — pre-filled with the real account numbers and amounts from your document, with an iterative refine loop ("shorter", "firmer")
- 🎭 Call rehearsal: practice the scary phone call — by voice — against an AI playing the organization's rep
- 💬 Grounded Q&A that quotes your document instead of giving generic advice
- 🌍 12 output languages — read a Japanese hospital bill in English, or a US lease in Chinese
History lives only in your browser's localStorage. Nothing sensitive is stored server-side.
How we built it
- Inference: Featherless AI's OpenAI-compatible API — Qwen3-VL-30B-A3B for vision analysis (structured JSON with normalized bounding boxes for each extracted fact) and Qwen2.5-72B for chat, drafting, and roleplay, with a multi-model fallback chain
- Backend: Python + FastAPI; Server-Sent Events for token streaming; PyMuPDF rasterizes PDF pages for the vision model; Pillow preprocesses uploads
- Frontend: dependency-free vanilla HTML/CSS/JS — camera-friendly on mobile; Web Speech API powers voice input and spoken rep replies; ICS calendar files generated client-side
- Deployed on Render; Dockerfile included for any container platform
Challenges we ran into
- Open-model reliability is the real project. Gated Llama weights, capacity exhaustion, and an upstream bug where the model emits pure
!token floods. The fix: a fallback chain across models, output sanitization, and flood detection that reroutes requests mid-flight - Structured output from prose models needed defensive JSON extraction — models wrap JSON in chatter, so we retry and parse around it
- A subtle async bug: blocking HTTP inside an async endpoint froze the whole server during vision calls — moved inference to a threadpool
- AI latency is a UX problem: first tokens can take 25s+. Streaming, elapsed timers, and honest "generating…" placeholders made it feel fine
- Vision grounding (bounding boxes for facts) required output validation and normalized coordinate mapping across resized images
Accomplishments that we're proud of
- The bbox grounding — hovering a fact flashes where it lives on the page — makes the AI's reading verifiable at a glance
- The scam check actually works: feed it a fake ticket demanding Bitcoin and it correctly flags "suspicious" with the right signals
- Drafts that pre-fill real numbers: the letter knows it's account RG-447821 and $685.13, not "Dear Sir/Madam [insert details]"
- The call rehearsal is the feature people remember — bureaucratic rep included
- It's a real, deployed product: live URL, 12 languages, privacy-respecting by design
What we learned
- Open-weights vision-language models are genuinely good enough for real document understanding — Qwen3-VL's grounding output worked with minimal coaxing
- Defensive engineering around model quirks (flood detection, template-kwarg side effects, JSON extraction) matters more than model choice
- In a hackathon, actionability beats accuracy benchmarks: a checklist + draft + rehearsed call is worth more than a slightly better summary
- Privacy can be a feature you ship in one line (localStorage) and a differentiator
What's next for Unfog
- Verified organization lookup (real phone numbers, official payment portals) to pair with the scam check
- Document diffing — "the new lease vs. your old one"
- Deadline push notifications; family-sharing view for helping parents decode their mail
- A PWA wrapper so it installs like a native camera-adjacent utility
Built With
- accessibility
- docker
- fastapi
- featherless-ai
- javascript
- python
- render
Log in or sign up for Devpost to join the conversation.