Inspiration
A teacher does not get stuck on the hard part of marking. She gets stuck on the pile around it: an incomplete sheet, a missing cell nobody noticed, then two write-ups of the same paper — talking points for the meeting, a calmer note for home.
That chase is hours. It is not judgement. It is operations.
Mark Map started from that pile. Not “an LLM that grades.” A class operations desk: catch the hole, refuse the lie, write both notes only when the row is true. Empty means missing, not zero. The parent never sees the teacher brief.
I wanted an agent that makes a professional faster at the work around the judgement — not an agent that takes the judgement.
What it does
Three doors, one class.
• Teacher loads Term 1 and Midterm, sees the chapter map, runs the desk, fills yellow cells, can grant or deny agent permissions. • Student sees their own numbers. No teacher brief. • Parent sees the same numbers. No teacher brief. No meeting language.
A Strands agent runs the desk. Tools, in a loop: incomplete rows → hotspots → try to write briefs → nag if a cell is empty. Python owns marks, completeness, and publish. A policy hook cancels write_mark, unlock_brief, and set_question_chapter before they run.
Briefs publish only when every cell is filled:
publish(row) ⟺ ∀ q ∈ Q: cell(row,q) ≠ ∅
Empty is not 0. “Give this student 8 on Q9” does not write a mark.
Picture reports are Tesseract + Python, not the model looking at pixels.
How we built it
• FastAPI + vanilla JS (three doors, cookie session). • One JSON store as the system of record: papers, rows, briefs, nags, permissions. • analyser.py / briefs.py / desk.py: deterministic kernel. The model does not recap a cycle Python already finished. • Strands Agents SDK: one desk Agent, four @tools, PolicyHook on BeforeToolCall. • Brains are swappable. Same tools: • scripted — published demo and tests (fast, honest). • LM Studio / oMLX — OpenAI-compatible /v1 (docs/lmstudio-ollama-mlx.md (https://github.com/Sankalp774/mark-map/blob/main/docs/lmstudio-ollama-mlx.md)). • Ollama — MLX-backed on Apple silicon (same doc). • Amazon Bedrock — BedrockModel Converse, instance role, no long-term API keys (docs/bedrock.md (https://github.com/Sankalp774/mark-map/blob/main/docs/bedrock.md)). • Eval gates: invented_marks = 0, leaked_teacher_briefs = 0.
Code (MIT): github.com/Sankalp774/mark-map (https://github.com/Sankalp774/mark-map)
Architecture: docs/architecture.svg (https://github.com/Sankalp774/mark-map/blob/main/docs/architecture.svg) — request path, multi-agent desk, Strands loop.
Challenges we ran into
Bedrock invoke is unauthorized on this AWS account. Models list. Converse returns ValidationException: Operation not allowed. get-foundation-model-availability returns authorizationStatus: NOT_AUTHORIZED for Nova, Llama, Mistral, and Claude, in several Regions, including as root. That is an account-level hold, not a missing IAM tick. Nova is not a Marketplace subscribe model. A Bedrock API key does not bypass it. A Support case was opened.
I did not fake a Bedrock trace. The Strands loop is real. The published demo uses scripted tools so judges see observe → plan → act without waiting on a blocked API.
Local LLMs (LM Studio, Ollama/MLX) run the same tools. Reasoning models were too slow for a five-minute video. Instruct 3B is usable; the desk is not a token window.
Two AWS identities (classic account vs a new-experience project) overwrote the CLI profile. Credits sit on the classic account. We stopped creating extra projects.
Prompt-only guardrails fail. “Do not unlock” in the prompt was not enough. The hook has to cancel the tool.
What we learned
• Professional agents should take the pile, not the judgement. • Python as system of record is easier to evaluate than asking the model to “be careful.” • Strands is the agent. Bedrock is only the brain. If the brain is blocked, keep the tools. • GetFoundationModelAvailability tells the truth faster than the playground. • A live demo that lies about its backend is worse than a scripted demo that names its tools.
What's next
When AWS sets authorizationStatus to AUTHORIZED: MARKMAP_DESK_MODEL=bedrock, App Runner instance role, EventBridge POST /api/desk/sweep. AgentCore is optional later, not a second product.
Until then, the desk already runs.
Built With
- agents
- amazon
- amazon-web-services
- app
- bedrock
- boto3
- css
- docker
- eventbridge
- fastapi
- html
- iam
- javascript
- json
- lm
- ocr
- ollama
- pytest
- python
- runner
- sdk
- strands
- studio
- tesseract
- uvicorn
Log in or sign up for Devpost to join the conversation.