Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for HomeworkHawk for Alexa+ — voice-first homework grading

What it does

"Alexa, check my kid's homework." The parent drops a finished math worksheet under the lamp. A photo is captured, and in 0.24 seconds — pure OpenCV 5 on CPU — the page is graded: per-question ✓/✗, digit-level error localization, and questions the system itself couldn't read confidently flagged for a human. Alexa+ then reads the results back by voice, names the questions worth reviewing together, and books a 20-minute review session for tomorrow evening, confirmed in conversation.

How we built it

This is a simulated Alexa+ experience (per the track's simulated-experience path) built as a web app around the real HomeworkHawk engine:

  • Voice front-end (alexa_sim/ in the repo): an Echo-Show-style device screen with a conversation thread and mic button. Voice out uses the browser's speechSynthesis, reading the same script an Alexa+ skill would — score summary, review list, escalated questions, follow-up booking.
  • Grading core: the production OpenCV 5 pipeline — quality gate (blur/glare), page rectification (IoU 0.998), illumination flattening, two-level print/handwriting separation, projection-profile segmentation, connected-component digit reading (93.3% exact-read). The Alexa+ front-end calls the exact same handle_photo entry point the web demo and the AWS Lambda deployment use.
  • Agentic layer (AWS Builder mini challenge): a Strands Agents SDK agent orchestrates five OpenCV tools — assess_photo, grade_page, re_read_question, escalate_to_parent, sheet_math. Vision evidence decides the next tool call: a re-binarization pass happens only when the first read's confidence is low, and truly unreadable questions escalate instead of guessing.
  • Open Source mini challenge: contributed the homeworkhawk-graded-worksheet-callback skill to CALLE-AI/awesome-phone-call-agents (PR #500) during the hackathon window — the phone-call variant of this same workflow.

Challenges we ran into

Voice-first reframing of a vision product. A dashboard can show a heat map; a voice interface has one shot at being understood. We rewrote the output as a spoken script — score first, review list second, escalated questions explicitly handed to the parent — and made "I couldn't read this one, please check by hand" a first-class utterance rather than a failure state.

Accomplishments

One engine, four surfaces: web demo, Lambda endpoint, phone call (CALL-E), and now voice (Alexa+ simulation). 0.24 s/page on CPU with zero ML-model dependencies, and a confidence-gated policy that never guesses about a child's work.

What we learned

The last mile of an agent is the interface a busy parent actually touches. The same grading result that dies in a dashboard becomes a conversation they can have while cooking.

What's next

Graduating the simulation toward a real Alexa+ Agent Skill / MCP server, and multi-child household profiles.

Track: Alexa+ · Mini challenges: AWS Builder (Strands SDK) + Open Source (PR #500)

Repo: github.com/Zaichek/homeworkhawk

Built With

  • strands-agents
Share this project:

Updates

Submission history