Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for HomeworkHawk — a Strands agent that grades homework
What it does
HomeworkHawk turns a single phone photo of a completed elementary math worksheet into a graded page — per-question ✓/✗, digit-level error localization, and a parent report — in 0.24 seconds, pure OpenCV 5 on CPU. A Strands Agents SDK agent orchestrates the whole job: every OpenCV stage is a callable tool, and the agent decides, from the visual evidence each tool returns, what happens next.
How we built it (Strands Agents at the core)
The agent (strands_agent/grader_agent.py in the repo) wires five tools into a Strands Agent with a strict grading policy:
| tool | what the agent sees | what it decides |
|---|---|---|
assess_photo |
blur/glare/darkness verdict | 'reshoot' → stop and tell the parent; 'accept' → grade |
grade_page |
per-question verdict, score, attempts, needs_attention list |
which questions deserve another look |
re_read_question |
Otsu re-read score for one question | accept the retry or escalate |
escalate_to_parent |
— | flag a question for a human — never guess about a child's work |
sheet_math |
independent arithmetic check | verify suspicious keys |
The policy is enforced by the system prompt and the tool results themselves: vision evidence changes the next tool call — a re-binarization pass is a second OpenCV call decided by what the first one saw. The deterministic vision core (unchanged across runs): quality gate → page rectification (IoU 0.998) → illumination flattening → two-level print/handwriting separation → projection-profile segmentation → connected-component digit reading (93.3% exact-read). Measured on 18 synthetic worksheets × 6 degradations, 180 questions; grading decision accuracy 85.6%, 0.15–0.30 s/page.
Challenges we ran into
Teaching an LLM agent restraint. The temptation for an agent loop is to keep calling tools until it likes the answer; our policy forces the opposite — after one re-read, a low-confidence question goes to the parent. Honest escalation had to beat confident guessing.
Accomplishments
The full loop runs end-to-end: photo → agent-orchestrated grading → parent report with per-question reasoning, every tool call grounded in vision output. E2E demo logs in the repo (strands_agent/demo.py).
What we learned
A well-tuned classical OpenCV pipeline plus a Strands agent as the decision layer is a genuinely useful product loop — no GPU, no fine-tuning, and the agentic part is exactly the part a plain pipeline can't do: knowing when to say "not sure".
What's next
Multi-line answers and fractions; a phone-call leg (built for the CALL-E track) that delivers results to the parent by voice.
Repo: github.com/Zaichek/homeworkhawk (MIT)
Built With
- strands-agents
Log in or sign up for Devpost to join the conversation.