DisasterDoc
The Problem
When identity or relief documents are damaged by floods, fire, water, tearing, or poor image quality, the challenge is not simply to "read" the document.
The real challenge is to determine what information can still be trusted.
A conventional OCR system may return a complete-looking string even when parts of a document are missing or unclear. An AI vision model may go one step further and infer what the missing information probably was.
For humanitarian and field-work scenarios, that can be dangerous.
DisasterDoc is built around a simple principle:
If the document doesn't show it, DisasterDoc doesn't claim it.
What We Built
DisasterDoc is an evidence-first damaged-document recovery system that processes degraded document images and separates visible evidence from inference.
Instead of producing a single "best guess", DisasterDoc classifies each important field as:
- RECOVERED — the available evidence sufficiently supports the value.
- PARTIAL — some evidence is visible, but the complete value cannot be safely established.
- UNRECOVERABLE — there is not enough usable evidence to recover the field.
For example:
| Field | Result |
|---|---|
| Name | RECOVERED |
| ID | RECOVERED |
| District | PARTIAL — "BENG..." |
| Date of Birth | UNRECOVERABLE |
When a field is PARTIAL or UNRECOVERABLE, DisasterDoc does not invent a completion.
How It Works
The system follows an evidence-preserving pipeline:
Damaged Document
↓
Image Preprocessing
↓
Memory-Safe OCR
↓
OCR Observations
(text + bounding box + OCR confidence)
↓
Deterministic Field Mapping
↓
Deterministic Classification
↓
RECOVERED / PARTIAL / UNRECOVERABLE
↓
Human Verification
↓
AI Commentary (Read-Only)
↓
JSON / PDF Evidence Report
Every recovered observation retains its connection to the original document evidence.
The system records:
- Raw OCR text
- Bounding-box coordinates
- OCR engine confidence
- Field mapping
- Evidence reasons
- Deterministic status
- Human-verification requirement
- Document SHA-256 hash
- Audit metadata
This makes the output auditable rather than simply generating a final answer.
Why Deterministic Classification?
The most important design decision was to separate evidence from explanation.
The deterministic pipeline decides what the document supports.
The AI layer does not modify:
- OCR text
- Bounding boxes
- OCR confidence
- Field status
- Evidence
- Recovered values
Instead, AI commentary is used only as a read-only explanation for cases where human verification is needed.
This creates a clear boundary:
Deterministic code decides facts. AI explains the evidence.
A multimodal model might infer that a damaged value is probably a particular name or date. DisasterDoc deliberately does not promote that inference into evidence.
What We Learned
Building DisasterDoc taught us that document recovery is not the same problem as document recognition.
A system can achieve impressive OCR output while still producing unsafe results when the source itself is incomplete.
We learned to think about:
- Evidence preservation — keeping the original OCR observations instead of returning only a final string.
- Uncertainty — allowing the system to explicitly say when information cannot be recovered.
- Human-in-the-loop workflows — treating uncertain results as verification tasks rather than failures to hide.
- AI boundaries — using generative AI where explanation is useful while preventing it from becoming the source of truth.
- Auditability — producing structured evidence that can be reviewed later.
Challenges We Faced
One of the biggest challenges was designing a pipeline that could handle damaged documents without turning missing information into hallucinated information.
OCR naturally produces noisy output when text is blurred, partially destroyed, rotated, or obscured. The difficult part was deciding when that output represented enough surviving evidence to call a field recovered.
We therefore focused on transparent field-level reasoning rather than a single overall prediction.
Another challenge was integrating AI without allowing it to silently alter deterministic results. We solved this by keeping AI commentary as a separate, read-only layer.
Built With
DisasterDoc was built using:
- Python
- Streamlit
- EasyOCR
- OpenCV
- Pillow
- NumPy
- Google Gemini API
- ReportLab
- Pytest
The application also uses SHA-256 hashing and structured JSON evidence to support traceability.
The Human Verification Principle
DisasterDoc is designed as an assistive evidence-recovery tool, not an authority that decides whether a person is entitled to a document, benefit, or service.
Its purpose is to reduce the amount of manual work required to inspect damaged documents while making uncertainty visible.
The final decision remains with a human verifier.
Recover what is visible. Flag what is uncertain. Never manufacture what is missing.
What Comes Next
Future versions could extend DisasterDoc with:
- More document formats and layouts
- Multilingual OCR
- Stronger field-level validation
- Additional image-damage simulations
- Better accessibility for field workers
- Offline deployment for low-connectivity environments
- More detailed evidence-review interfaces
The long-term goal is to make damaged-document recovery more transparent, auditable, and human-verifiable rather than simply more automated.
Built With
- ai
- analysis
- artificial
- computer
- data
- document
- easyocr
- gemini
- generative
- human-in-the-loop
- intelligence
- learning
- machine
- ocr
- opencv
- processing
- python
- streamlit
- vision
Log in or sign up for Devpost to join the conversation.