Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for HomeworkHawk for Alexa+ — voice-first homework grading
What it does
"Alexa, check my kid's homework." The parent drops a finished math worksheet under the lamp. A photo is captured, and in 0.24 seconds — pure OpenCV 5 on CPU — the page is graded: per-question ✓/✗, digit-level error localization, and questions the system itself couldn't read confidently flagged for a human. Alexa+ then reads the results back by voice, names the questions worth reviewing together, and books a 20-minute review session for tomorrow evening, confirmed in conversation.
How we built it
This is a simulated Alexa+ experience (per the track's simulated-experience path) built as a web app around the real HomeworkHawk engine:
- Voice front-end (
alexa_sim/in the repo): an Echo-Show-style device screen with a conversation thread and mic button. Voice out uses the browser's speechSynthesis, reading the same script an Alexa+ skill would — score summary, review list, escalated questions, follow-up booking. - Grading core: the production OpenCV 5 pipeline — quality gate (blur/glare), page rectification (IoU 0.998), illumination flattening, two-level print/handwriting separation, projection-profile segmentation, connected-component digit reading (93.3% exact-read). The Alexa+ front-end calls the exact same
handle_photoentry point the web demo and the AWS Lambda deployment use. - Agentic layer (AWS Builder mini challenge): a Strands Agents SDK agent orchestrates five OpenCV tools —
assess_photo,grade_page,re_read_question,escalate_to_parent,sheet_math. Vision evidence decides the next tool call: a re-binarization pass happens only when the first read's confidence is low, and truly unreadable questions escalate instead of guessing. - Open Source mini challenge: contributed the
homeworkhawk-graded-worksheet-callbackskill to CALLE-AI/awesome-phone-call-agents (PR #500) during the hackathon window — the phone-call variant of this same workflow.
Challenges we ran into
Voice-first reframing of a vision product. A dashboard can show a heat map; a voice interface has one shot at being understood. We rewrote the output as a spoken script — score first, review list second, escalated questions explicitly handed to the parent — and made "I couldn't read this one, please check by hand" a first-class utterance rather than a failure state.
Accomplishments
One engine, four surfaces: web demo, Lambda endpoint, phone call (CALL-E), and now voice (Alexa+ simulation). 0.24 s/page on CPU with zero ML-model dependencies, and a confidence-gated policy that never guesses about a child's work.
What we learned
The last mile of an agent is the interface a busy parent actually touches. The same grading result that dies in a dashboard becomes a conversation they can have while cooking.
What's next
Graduating the simulation toward a real Alexa+ Agent Skill / MCP server, and multi-child household profiles.
Track: Alexa+ · Mini challenges: AWS Builder (Strands SDK) + Open Source (PR #500)
Repo: github.com/Zaichek/homeworkhawk
Built With
- strands-agents
Log in or sign up for Devpost to join the conversation.