Inspiration
Iu Mien is an endangered language with few digital resources. Most of its vocabulary lives in scanned print dictionaries: hard to search, full of OCR errors, and under copyright. Community members can't easily look up a word, hear it, or see how it's used. I wanted to unlock that knowledge for research without violating authors' rights or the community's control over its own language data.
What it does
- Extracts scanned dictionaries into structured data: headwords, homograph numbers, parts of speech, numbered senses, phrases nested under the right sense, examples with translations, cultural notes, synonyms, antonyms, variants, contractions and classifiers.
- Repairs OCR errors using the language itself. Every Iu Mien syllable has the shape $\sigma = C_{\text{initial}} + R_{\text{rhyme}} + T_{\text{tone}}$, so impossible syllables are flagged and fixed. The dictionary is alphabetical, so misread headwords show up as out of order. Cross-references must point to real entries. A local vision model re-reads doubtful headwords from the page image, with a constrained multiple-choice mode built for homograph numbers.
- Merges many sources into one "Dictionary Universe" that links the same word across dictionaries, with support for dialects, audio, other scripts (Thai, Lao) and a time-aligned corpus.
- Enforces licensing and community governance in code: the public sees only open material, community members see community material, researchers see reference-only sources. A facts-only export strips copyrighted expression before anything reaches the public site.
- Serves it through a bilingual website: Iu Mien and English search, tone-insensitive search, entries with sense-grouped phrases, audio and cross-source links, plus a JSON API.
How I built it
- Extraction: Python + pdfplumber, reading font and position data from the scan's text layer; gutter detection, skew-tolerant indentation to tell headwords from phrases, and section-heading detection (including multi-letter initials like Nz – nz).
- Correction: rule-based OCR repair (subscripts read as letters, c/e, rn/m, stray capitals), a data-driven syllable inventory, an alphabetical-order check (longest ordered run) that follows the dictionary's own tone-aware sort order, and cross-reference validation.
- Vision: Qwen3-VL-8B running fully on-device with MLX, so no page ever left the laptop. It started with free-form reading; after measuring its errors, I built a constrained multiple-choice mode.
- Human in the loop: review sheets (Correct / Wrong / type the right spelling) imported back into the database. Every automatic fix is logged, and every human decision survives re-runs.
- Lexicon and site: SQLite → Django (models, admin, licence-gated views, JSON API). The navigation is an embroidered "story band" inspired by Iu Mien textiles.
Challenges I ran into
- OCR errors that produce other real words: a subscript ₄ read as "m" turned aa₄ into am, itself a valid word. Only alphabetical order and page-image checks caught these.
- A general vision model got letters wrong: it got 47 of 54 letter changes wrong, mostly reading a subscript ₁ as the letter "i". Human review exposed this, so I stopped auto-applying letter changes and built a multiple-choice mode instead.
- Layout traps: scan skew, off-centre columns, and multi-letter section headings that had silently hidden whole pages. Fixing the headings recovered 1,200+ entries.
- Copyright: building a system that can use a copyrighted source for research while guaranteeing it never appears publicly.
Accomplishments that I'm proud of
- 30,000+ entries structured from ~790 scanned pages, with 108,000+ phrase-to-word links, 1,800+ cultural notes and 5,500+ example sentences (internal research).
- Thousands of OCR errors corrected, with a person reviewing every uncertain case.
- Licence enforcement that fails safe: the public import refuses any non-open data.
- All AI processing ran locally, consistent with community data sovereignty.
What I learned
- A language's own structure (its syllables, tones and sort order) is a powerful error-correction signal.
- General vision models need constrained questions and measured accuracy before they can be trusted.
- Responsible language technology is as much governance as code.
What's next for Dictionary Universe
- Community review of entries and cultural notes by Iu Mien speakers.
- More sources: community wordlists and speaker recordings.
- Running the multiple-choice homograph check across the full dictionary.
- AI-assisted semantic and ethnographic tagging.
- A translation model trained only on appropriately licensed data.
- Licensing conversations with the rights holders of legacy dictionaries.
- Oral history integration: connecting recorded elder narratives to dictionary entries.
- A replicable model: documenting the methods so scholars and communities can build the same kind of platform for other endangered languages.
Built during &hacks vs. before
- Built during the event (12 PM Sat to 12 PM Sun): extraction pipeline, OCR repair and validation, local vision checks, review tooling, the Dictionary Universe lexicon, facts-only export, licence-gated Django dictionary app and front end.
- Existed before: the host website shell and an earlier prototype.
- Data and IP: no third-party dictionary content is included in this submission or its repository; the source dictionary was processed for internal research only.
- AI tools: Claude (coding assistant), Qwen3-VL (local vision model).
Log in or sign up for Devpost to join the conversation.