Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for HomeworkHawk — an agentic vision homework grader
What it does
HomeworkHawk turns a single phone photo of a completed elementary math worksheet into a graded page: per-question ✓/✗, digit-level error localization ("the 3rd digit should be 8"), and a heat-map overlay — in 0.24 seconds, pure OpenCV 5 on CPU.
Why it is agentic, not just a pipeline
Every OpenCV 5 stage is a callable tool and a decision policy sits on top. Vision evidence changes the next action:
- blur / glare metrics fail the quality gate → the system asks the parent to re-shoot
- page quadrilateral not found → re-run detection with looser Canny parameters
- per-digit confidence < 0.72 → re-binarize the row (Otsu) and re-read — a second OpenCV call decided by what the first one saw
- still uncertain → escalate that question to the parent — the system says "not sure" instead of guessing about a child's work
Every decision lands in a replayable trace (tool, parameters, perception, decision) — see evidence/*/result.json in the repo.
How we built it (OpenCV 5 core)
- Quality gate — Laplacian-variance blur, clipped-white glare, darkness ratios
- Page rectification — Canny → contours → approxPolyDP quad → getPerspectiveTransform/warpPerspective (IoU 0.998)
- Illumination flattening — background-estimate division (kills phone shadows)
- Print/handwriting separation — two-level intensity classification (toner ≪ pencil) + dilated-print subtraction
- Question segmentation — projection profiles, ruling-line exclusion, question-number header filter, answer zone right of the "="
- Answer reading — connected-component glyphs, fragment-only merging, aspect-preserved fixed-grid cosine vs digit-template banks
- Agentic loop — per-digit confidence gates re-binarize / re-read / escalate
AWS: the identical handler (shared byte-for-byte) runs on Lambda behind a Function URL with an opencv-python-headless 5 layer (Graviton or x86_64); S3 for per-stage evidence with lifecycle expiry; CloudWatch for per-stage metrics. Free-tier sized: 0.15–0.30 s CPU per page.
Measured results
18 labeled synthetic worksheets × 6 degradations (perspective warp, diagonal shadow, blur, faint pencil, half-erased answer) — 180 questions, 30% wrong answers injected:
| Metric | Result |
|---|---|
| Page-detection IoU | 0.998 |
| Question segmentation | ≈100% clean/shadow/faint/erased; 81.1% overall |
| Handwritten answer exact-read | 93.3% |
| Grading decision accuracy | 85.6% |
| Blur photos correctly rejected | 3/3 |
| Latency | 0.15–0.30 s/page, CPU |
Published failure cases
Perspective-warped edges lose 1–2 rows; the template reader is font-biased (unusual handwriting escalates rather than errors silently); half-erased pencil ghosts sometimes still read. All runs including failures are in evidence/.
Challenges we ran into
Separating pencil from toner under household shadows (solved by flatten-then-two-level-intensity); not clipping taller-than-print handwriting during row segmentation; and honest escalation design — making the system comfortable saying "not sure".
Accomplishments we're proud of
0.24 s/page end-to-end on CPU with zero ML-model dependencies, and an agentic loop where every OpenCV re-analysis pass is logged and replayable.
What we learned
That a well-tuned classical OpenCV pipeline plus a confidence-gated action policy can deliver a genuinely useful product loop without any GPU — and that publishing failure cases is a feature.
What's next for HomeworkHawk
Multi-line answers, fractions, drawn figures; a small on-device digit classifier to augment the template bank (traces already record the training crops); the COOL/Graviton benchmark path.
Responsible use
Transient processing, no faces required, evidence expires, human approval gates every ambiguous verdict, and the product grades + explains — it never solves the exercise for the child.
Log in or sign up for Devpost to join the conversation.