Inspiration

Kenya's Competency-Based Curriculum asks teachers to do more than ever — plan lessons around specific KICD strands and sub-strands, keep schemes of work aligned to the term calendar, and produce evidence-based post-lesson reflections — on top of actually teaching. None of that admin work is optional, and almost none of it needs to consume as much time as it does.

We also saw a specific, less obvious risk: the easy version of "AI for lesson planning" is a model that sounds fluent and confident while quietly inventing a learning outcome or activity that was never in the curriculum. When the output is what children are taught, that's not a minor bug — it's the whole problem with applying generic AI here carelessly. We wanted an agent that treats the real KICD curriculum as the one source of truth it's never allowed to contradict.

What it does

The agent grounds the full teacher planning workflow in the official KICD curriculum, end to end:

  • Curriculum Explorer — real semantic and filtered search over ingested KICD Grade 4–6 curriculum designs, returned as evidence cards with exact source-page citations, not paraphrased summaries.
  • Term plan generation — a teacher selects real curriculum evidence for one sub-strand, and the agent organizes it into weekly scheme-of-work rows, paced by the sub-strand's actual lesson count from the curriculum (a 7-lesson sub-strand becomes three real weeks, not one arbitrary block), following the current, officially-used "rationalized" scheme format rather than an idealized structure.
  • Daily lesson plan generation — grounded in a specific term-plan row's own content, producing an introduction, activity sequence, and assessment section drawn only from what was actually retrieved.
  • Post-lesson reflection — the teacher records their own evidence, and the agent summarizes it back to them in plain language, with achievement-status language stripped in code, not just discouraged in a prompt — the decision of whether an outcome was met stays with the teacher, always.
  • A genuinely agentic assistant — built on the Strands Agents SDK, with a curriculum-search tool it decides for itself whether to call, depending on whether the context it already has is enough to answer.
  • Word document export and a real confirmed-work Library, backed by PostgreSQL, so nothing becomes an official record without an explicit teacher confirmation step.
  • Dual model-provider support — Google Gemini and AWS Bedrock (Qwen and Amazon Titan), swappable via configuration, so the same grounded generation logic runs on either provider.

How we built it

  • Strands Agents SDK for orchestration — the deterministic generation functions (term plans, daily lessons, reflection summaries) run through Strands' structured-output interface; a separate, genuinely autonomous Strands Agent with a curriculum-search tool powers the free-text assistant, where tool-selection ambiguity is actually appropriate rather than a risk to avoid.
  • ChromaDB for the curriculum knowledge base, ingested from real KICD Grade 4–6 curriculum design PDFs via Docling table extraction, with metadata (grade, subject, strand, sub-strand, real lesson counts) used directly to pace generated schemes correctly.
  • PostgreSQL for all application state — users, schemes of work, lesson plans, evaluation records, and an audit trail of every confirmation — with generation output always staying an unconfirmed draft until a teacher explicitly reviews and confirms it.
  • FastAPI backend, Next.js/React frontend, python-docx for Word document generation.
  • Deliberate, code-level grounding enforcement, not just prompt instructions: a generated field with no supplying evidence is forced empty in code, regardless of what the model returns; a reflection summary that judges achievement status has that sentence stripped programmatically before it ever reaches the teacher.

Challenges we ran into

  • The KICD data pipeline broke before we could even start generating anything. Our first extraction source — a scraped Google Drive preview of the official curriculum designs — had destroyed all table structure; a broken source page has no coordinate information left for any parser to recover, no matter how well-written. The real fix wasn't better parsing, it was a better source: we found and verified legitimate, freely downloadable mirrors of the actual KICD PDFs, subject by subject.
  • Real teacher practice was more informative than the curriculum design document alone. Reviewing six real, currently-used schemes of work revealed a "rationalized" convention — identical grounded content repeated across a sub-strand's full lesson allocation — that we hadn't anticipated, and which turned out to be a safer generation target than inventing distinct day-to-day variation would have been.
  • Gemini refuses to generate text too close to its source material. Grounding a model tightly in copyrighted curriculum text runs directly into the same safety filter designed to catch near-verbatim reproduction — an ironic collision between "be maximally faithful" and "don't just copy." The fix was a retry that explicitly asks for rewording while staying grounded, with the specific refusal reason preserved through a small Strands model subclass, since the SDK's default behavior discards that detail.

What we learned

  • Grounding an agent in a real, authoritative dataset changes the whole shape of the problem, from "generate something plausible" to "retrieve and present the right thing well" — most of the hard work lives in retrieval and data quality, not prompt cleverness.
  • Prompt instructions alone are not a reliable enforcement mechanism for anything that actually matters. Every hard constraint in this project that we cared about — never inventing missing content, never suggesting an achievement judgment, even plain-text formatting — needed a code-level check behind it, because the model does not reliably follow the prompt alone.
  • Silent infrastructure failures can look identical to logic bugs in your own code. The rogue Postgres installation and the reverted merge both wasted real time specifically because the symptom looked like "our fix isn't working," when the actual problem was one layer further down.

What's next

  • The original hackathon framing called for a fully autonomous background agent — proactively drafting and nudging teachers via WhatsApp rather than a workspace they open themselves. We deliberately prioritized building a complete, deeply evidence-grounded application first; a background/notification layer on top of the same generation and grounding logic is the natural next phase, not a redesign.
  • Expand curriculum coverage beyond the current pilot subjects and grades. Bedrock Knowledge Bases Integration: Migrating local ChromaDB ingestion directly to Amazon Bedrock Knowledge Bases with OpenSearch Serverless for automated sync with official Ministry of Education syllabi. -Multimodal Resource Extraction with Qwen VL: Leveraging Qwen3 VL 235B A22B on Bedrock to allow teachers to photograph physical textbook exercises and automatically extract them into supplementary lesson activities. -Multi-Tenant Teacher Auth: Rolling out Amazon Cognito for secure school-level and county-level authentication.

Built With

Share this project:

Updates

Submission history