Inspiration

Everyone has received a piece of mail they couldn't fully understand: a medical bill full of CPT codes, a parking ticket with escalating fines buried in legalese, a lease notice hiding an auto-renewal trap. The real cost isn't confusion — it's missed deadlines, overpaid fees, and unchallenged denials. The burden falls hardest on students, immigrants, and anyone reading in a second language.

Generic chatbots can define jargon, but they can't read your document, find your deadline, or draft your reply. We wanted to close that last mile — from "what does this say?" to "it's handled."

What it does

Unfog is a paperwork decoder. Snap a photo (or drop a PDF, or paste text) and get:

  • 📖 A plain-language summary + one-line ELI5 of what the document actually means
  • 🔑 Key facts extracted — and visually grounded: hover any fact and it highlights where that text sits on your document
  • ⏰ Deadlines with countdowns and one-tap calendar (.ics) export
  • 🛡️ A legitimacy check — it flags scam patterns (gift-card demands, arrest threats, fake URLs) before you pay a fake ticket
  • ✅ An actionable checklist where every step has a "how do I do this?" button
  • ✍️ Ready-to-send drafts — dispute letters, phone scripts, questions to ask — pre-filled with the real account numbers and amounts from your document, with an iterative refine loop ("shorter", "firmer")
  • 🎭 Call rehearsal: practice the scary phone call — by voice — against an AI playing the organization's rep
  • 💬 Grounded Q&A that quotes your document instead of giving generic advice
  • 🌍 12 output languages — read a Japanese hospital bill in English, or a US lease in Chinese

History lives only in your browser's localStorage. Nothing sensitive is stored server-side.

How we built it

  • Inference: Featherless AI's OpenAI-compatible API — Qwen3-VL-30B-A3B for vision analysis (structured JSON with normalized bounding boxes for each extracted fact) and Qwen2.5-72B for chat, drafting, and roleplay, with a multi-model fallback chain
  • Backend: Python + FastAPI; Server-Sent Events for token streaming; PyMuPDF rasterizes PDF pages for the vision model; Pillow preprocesses uploads
  • Frontend: dependency-free vanilla HTML/CSS/JS — camera-friendly on mobile; Web Speech API powers voice input and spoken rep replies; ICS calendar files generated client-side
  • Deployed on Render; Dockerfile included for any container platform

Challenges we ran into

  • Open-model reliability is the real project. Gated Llama weights, capacity exhaustion, and an upstream bug where the model emits pure ! token floods. The fix: a fallback chain across models, output sanitization, and flood detection that reroutes requests mid-flight
  • Structured output from prose models needed defensive JSON extraction — models wrap JSON in chatter, so we retry and parse around it
  • A subtle async bug: blocking HTTP inside an async endpoint froze the whole server during vision calls — moved inference to a threadpool
  • AI latency is a UX problem: first tokens can take 25s+. Streaming, elapsed timers, and honest "generating…" placeholders made it feel fine
  • Vision grounding (bounding boxes for facts) required output validation and normalized coordinate mapping across resized images

Accomplishments that we're proud of

  • The bbox grounding — hovering a fact flashes where it lives on the page — makes the AI's reading verifiable at a glance
  • The scam check actually works: feed it a fake ticket demanding Bitcoin and it correctly flags "suspicious" with the right signals
  • Drafts that pre-fill real numbers: the letter knows it's account RG-447821 and $685.13, not "Dear Sir/Madam [insert details]"
  • The call rehearsal is the feature people remember — bureaucratic rep included
  • It's a real, deployed product: live URL, 12 languages, privacy-respecting by design

What we learned

  • Open-weights vision-language models are genuinely good enough for real document understanding — Qwen3-VL's grounding output worked with minimal coaxing
  • Defensive engineering around model quirks (flood detection, template-kwarg side effects, JSON extraction) matters more than model choice
  • In a hackathon, actionability beats accuracy benchmarks: a checklist + draft + rehearsed call is worth more than a slightly better summary
  • Privacy can be a feature you ship in one line (localStorage) and a differentiator

What's next for Unfog

  • Verified organization lookup (real phone numbers, official payment portals) to pair with the scam check
  • Document diffing — "the new lease vs. your old one"
  • Deadline push notifications; family-sharing view for helping parents decode their mail
  • A PWA wrapper so it installs like a native camera-adjacent utility

Built With

Share this project:

Updates

Submission history