Inspiration
Worksheets fail for a boring reason: they don't fit. Too long for the period, pitched above the class's reading level, all recall questions and no application, missing half the standard it's supposed to cover. A teacher usually finds this out mid-class, not before printing.
I wanted to build something that catches that before the worksheet goes out, not another tool that generates content, but one that keeps checking whether the content actually works, the whole time you're building it. WebMCP made the second half of that idea possible: instead of an AI suggesting text in a chat window, the page itself could hand an agent real, callable tools (check_reading_level, check_time_estimate) and let it act on the actual worksheet directly, alongside the teacher, not instead of them.
What it does
Sheshat is a worksheet workspace where a teacher and an AI agent build the same document together, live. You start from a photo of an old worksheet, a whiteboard, or pasted notes, and the app drafts real, grounded questions from what's actually there. From then on, four checks stay visible at all times: completion time, reading level, question-type balance, and standards coverage.
You can edit the worksheet directly (rewrite, reorder, delete a question by hand) or tell the agent what to change in plain language ("swap the fractions question for something on decimals"), and it calls the exact same tools a direct edit would trigger. Either path updates the same board and reruns all four checks immediately. The worksheet isn't "done" when it has content, it's done when all four checks pass.
A dashboard sits in front of the editor: the agent can list your worksheets, filter them, create a new one, and open it, not just edit inside one page, but navigate the whole workspace. In total, the app exposes 24 WebMCP tools across two route-scoped surfaces: 4 on the dashboard, 20 inside the editor.
How we built it
React + Vite on the frontend, a small Node API with SQLite for persistence, deployed on Vercel.
Every constraint check is real, deterministic logic, not a model guess: word/question-count heuristics for time, sentence-complexity heuristics for reading level, straightforward tallying for question mix and standards coverage. The kind of thing that has to stay reliable even if generation has a bad day.
WebMCP sits on top of all of it via document.modelContext.registerTool(). Tools are scoped by route: dashboard tools when you're browsing your worksheets, the full editing toolset once you're inside one, so an agent only sees the actions that are actually relevant to where the teacher is.
Challenges we ran into
The hardest bug wasn't glamorous: an early version of swap_question wrote its own input instruction into the question text instead of a real generated question. The worksheet would silently fill with prompt text instead of usable content. Fixed with real validation guards: reject output that's empty, malformed, or suspiciously identical to the input, and fail visibly instead of committing broken content.
A security review also surfaced two real issues: the local API would accept state-changing requests from any website open in another tab (no origin validation), and the image-source tool would happily fetch and render an arbitrary remote URL. Both are fixed now, with strict origin/content-type checks on every mutation endpoint, and image input restricted to inline data only.
Accomplishments that we're proud of
The four-check system actually does what it claims. I can point to a real before/after: a generated draft opening at 2 of 4 checks satisfied, one direct edit and one agent command later, 4 of 4. Not a mocked demo state, the same recalculation runs every time, from real logic, on real edits.
I'm also proud of catching my own bugs before they became someone else's problem (the swap bug, the security findings) by actually testing the exact flows a judge or user would try, not just trusting that generated code worked.
What we learned
That "AI-powered" is doing a lot of work to hide whether something actually functions. It was tempting more than once to let a claim stand ("it's fixed," "it's dynamic") without re-testing it myself. The moments this project actually got better were the ones where I stopped and checked instead.
I also learned the real difference between a chatbot suggesting worksheet text and an agent calling structured tools on a live page: the former can't reliably track question IDs, ordering, or state across a conversation. WebMCP's whole value is that the page, not the chat history, is the source of truth.
What's next for Sheshat
Real teacher testing. Everything here has been verified by me, not yet by the people it's actually for. Beyond the hackathon scope: accounts, more subjects and grade bands beyond the current Grade 4 to 7 math focus, and eventually making the core "does this fit my class" checking available without requiring a WebMCP-enabled browser, so any teacher can use it, not just ones testing experimental browser flags.
Built With
- apifirebase
- gemini
- node.js
- react
- sqlite
- typescript
- vercel
- vite
- webmcp
Log in or sign up for Devpost to join the conversation.