Inspiration

Helping a child practice reading shouldn't mean adding another app to check every day. Most tools for this ask a parent, or a busy tutor, to log in, pick an assignment, mark it, and remember to send the next one — one more thing competing for attention in an already full day.

The Everyday Agents track asks for an agent that takes the busywork out of daily life and runs quietly in the background, only surfacing when there's a real decision to make. That framing matched something I'd actually want to build: a reading tutor that lives entirely in email. No dashboard, no login, no app icon. You enroll a student once, and from then on the only thing anyone ever sees is an email — a worksheet arriving, a reply going out, and eventually a "you've finished" note. Everything in between happens on its own.

What it does

Tutor is a background AI tutor for Year 6 reading comprehension, built with the Strands Agents SDK and deployed on Amazon Bedrock AgentCore Runtime.

  • Enroll a student (one DynamoDB write) and Tutor generates a fresh reading passage and question set with Amazon Bedrock, renders it to a PDF, and emails it out.
  • The student photographs their handwritten answers and replies to the email.
  • Tutor reads the handwriting, grades it against the worksheet's own answer key, and decides what happens next: a new worksheet on a different topic if the student passed, another attempt if they didn't, or, after five passes , a congratulations email instead of another assignment.
  • If a photo is too blurry to grade reliably, Tutor says so and waits, rather than guessing.

Nothing in that loop is scripted from the outside. Every step - generating, grading, deciding what's next, sending - is a tool the agent calls itself.

How we built it

The system is two deployed pieces working together:

  • TutorAgent : the Strands agent itself, running as a Container-build agent on AgentCore Runtime. It has two flows : generate and send (triggered on enrollment) and grade a submission (triggered on an inbound reply). Tools include worksheet generation, PDF rendering (Puppeteer), email sending with SES, DynamoDB state read/write, and handwriting grading via Bedrock's multimodal input.
  • tutor-infra : a separate CDK stack providing the automatic triggers: a DynamoDB table with Streams enabled for enrollment, an SES receipt rule and S3 bucket for inbound replies, and two Lambdas that invoke the agent when either event fires.

I started with a local, no-AWS test harness to prove the core logic - generation quality, PDF rendering, and grading accuracy against real handwritten photos - before wiring any of it into cloud infrastructure. That local-first approach paid off: by the time I was debugging deployment issues, I already knew the agent logic worked, so every new failure was clearly an infrastructure problem, not a logic one.

Challenges we ran into

  • A tool-call argument that silently hung for two to three minutes. render_worksheet_pdf originally returned a PDF as base64 for the next tool to consume, and every request would hang with no error. The cause: with Strands, tool-call arguments are generated by the model itself, token by token, so a PDF-sized base64 string meant the model was silently retyping the entire PDF before the next tool could run. Fixed by returning a file path instead and reading the bytes from disk inside the tool - a lesson specific to how agentic tool-calling actually works, not something that would ever surface in a normal API integration.
  • AgentCore's environmentVariables config field silently did nothing. Setting it in agentcore.json and redeploying produced no error, but the value never reached the running container, confirmed by querying the deployed runtime directly and getting null back every time. Root-causing it meant tracing the actual deploy path through the CLI's generated CDK into the @aws/agentcore-cdk construct library, where I found the construct that correctly wires environment variables through is only ever instantiated for MCP server runtimes, never for an application agent's runtime — and separately, the schema's real field name is envVars (an array of {name, value} pairs), not environmentVariables (an object), which the config validator accepted anyway without complaint.
  • Model trying to start a conversation when there is none After the initial deploy to AgentCore, the Claude model started asking questions when prompted by the tool instead of generating a worksheet. I had to remove an MCP server added as tool by default by the AgentCore CLI bootstrap and adjust the prompt to explicitly instruct the model not to open a conversation when prompted. It was a lesson on how models need continuous training and why AI/ML ops is a necessary piece in unit testing the tools that the agents use as we develop more of them.

Accomplishments that we're proud of

Getting the full loop working automatically, end to end, with real email: writing one record to DynamoDB is the only manual step in the entire system. Everything after that - generating a worksheet, sending it, reading a photographed reply, grading real handwriting, deciding what happens next, and eventually sending a congratulations email - runs without anyone touching a keyboard again.

I'm also proud of actually tracking down the harder bugs to their real root cause rather than settling for a workaround that merely seemed to work. The environmentVariables config issue is a good example: instead of accepting "it doesn't propagate, so hardcode a default," I traced the actual deploy path through the CLI's generated CDK into the @aws/agentcore-cdk construct library and found the real cause - the construct that wires environment variables through is only ever instantiated for MCP server runtimes, and separately, the schema's real field name is envVars (an array), not environmentVariables (an object), which the config validator silently accepted anyway. Finding and fixing the actual bug, not just working around it, is the difference between a demo and something I'd trust to keep running.

What we learned

That a model's own output can be the bottleneck, not just its reasoning. Any data a tool hands back that has to flow into a later tool call gets regenerated token by token by the model in between, so passing anything large (a PDF, an image) as a return value is a real performance trap, not just a style choice. File paths and references, not payloads, are the right shape for that boundary.

That AWS's own construct libraries can have real gaps, and the only way to find them with confidence is to read the actual source the CLI generates, not just the documented config schema. The agentcore.json schema accepted a field that the deploy path never read at all — a silent, well-hidden gap that only showed up by tracing code, not by re-reading docs.

That debugging infrastructure benefits from the same discipline as debugging code: isolate variables one at a time (raw CLI vs. CDK, one region vs. another), and trust what you can directly observe (the actual deployed bytes, the actual synthesized template) over what should logically be true.

What's next for Tutor

  • Extend beyond reading comprehension to writing assignments, and support more than one year level.
  • Replace the manual DynamoDB enrollment step with a simple enrollment form for parents and tutors - a small web form (API Gateway + Lambda, writing to the same TutorStudents table) so adding a student doesn't require the command line at all.
  • Automated progress reports for teachers. Since every submission is already graded and stored - pass/fail, topic, score - a scheduled Lambda could roll this up per student, or across a teacher's whole class, into a periodic summary email: topics covered, pass rate, questions students are consistently missing. Kept in the same spirit as the rest of the system - no dashboard to check, just a report that shows up in an inbox on its own schedule.
  • Guardrails for the Agent to ensure that the input and output from the Agent is appropriate content wise.
  • AI Ops to maintain the standards of the agent during continuous development and deployments.

Built With

Share this project:

Updates

Submission history