-
-
-
The race post by post: checkpoints coloured by staffing, the schedule by site, and the chat with the agent.
-
A shift one person short: why nobody else qualifies, and who is closest to qualifying.
-
Race day: two no-shows announced at 11:00, the plan repaired with the fewest moves, and a question back to the organizer.
Inspiration
Last September I ran the SwissPeaks Marathon: 46 km and 2,483 m of climb in the Swiss Alps, eight and a half hours. I remember the volunteers as much as the trail. Someone had put each of them at the right place at the right hour, and that someone works with a roadbook PDF made of images, a sign-up form where a hundred people write whatever they want, and a WhatsApp group, on their evenings. Big races buy a platform. Small races have Excel and goodwill. Rubalise is for them. The name is the striped tape that marks the trail.
What it does
The organizer drops the roadbook and the sign-up form in a chat. The agent reads the PDF pages as images and returns a closed race sheet (checkpoints, first and last runner, cut-offs) with a list of what it was not sure about. It translates every free-text line of the form into fields a solver understands, keeping the remarks tagged. A template turns the race into 50 shifts. A constraint solver assigns 99 volunteers under hard rules (availability, night refusals, minors, skills, travel between sites) and weighted goals (minimum staffing, preferences, pairs, workload). A deterministic checker lists violations. A second agent, the contradictor, rereads the plan with a blank context and reports what no rule describes: a post leader whose three unlisted companions carry the whole post, a thirteen-hour day.
When someone is missing, the agent does not stop at "no". It names who is closest and what to change: three people are free and licensed, they only lack a 4x4.
Nothing reaches the volunteers unless the organizer explicitly confirms. Publishing, waiving a rule and changing a rule go through a gate that lives in the code, not in the prompt. On race day, the published plan is a commitment: every change has a cost, higher for people already on site, and the solver repairs with the fewest moves. Finished shifts never move.
A web cockpit shows the race post by post: a timeline of checkpoints coloured by staffing, the schedule by site, the detail of a shift with its team and its gaps, and the chat with the agent.
How we built it
Two Strands agents: the orchestrator (Claude Sonnet 4.6) holds the conversation and calls the tools, in a known order during preparation and freely on race day; the contradictor (Claude Opus 4.6) only reads and reports. The tools are exposed by an MCP server (16 tools: read roadbook, translate form, build shifts, solve, verify, rules, derogations, journal, publish) and are testable without any model. The solver is OR-Tools CP-SAT. Everything the agents know lives in a JSON state per session, not in their memory.
Hosting: Amazon Bedrock AgentCore Runtime (CodeZip, Python 3.12), the MCP server running inside the session. The cockpit is a single HTML page behind a Lambda relay and an HTTP API; because a plan takes longer than the API's 30-second limit, the relay starts the job, stores the result in S3 and the page polls.
Challenges we ran into
Making the language model useful without letting it plan: it is bad at counting and at holding fifty constraints at once, so every model step fills a closed JSON schema that code validates, and every step is measured against ground truth. Keeping a published plan stable on race day, which meant a second objective in the solver with change costs. Making the human gate real when the agent is hosted: it executes only if the organizer's own message contains an explicit confirmation. And the usual: Bedrock quotas, a blocked Lambda function URL, a 30-second API limit, sessions that sleep.
Accomplishments that we're proud of
Form translation: 99 out of 100 lines right, and the one difference is the runner whose Saturday the model correctly removed. Roadbook: all six checkpoints exact. A full plan in ten seconds with the reason for every gap and the nearest candidates named. Race day: five no-shows repaired in half a second with one removal and six additions. The contradictor finding, unprompted, the post held up by one family.
What we learned
A solver plus a language model that reads and explains beats a language model that plans. A second agent with a blank context catches what the first one cannot see, because it did not build the plan. And a guard in a prompt is not a guard; a guard in the code is.
What's next
Volunteer briefs written by the model for each person, a post-race review that reads the journal and proposes next year's rules, durable state across sessions, and other event templates (a triathlon is another template file, not another program).
Blog posts
Three posts on the AWS Builder Center tell the build story:
Built With
- amazon-api-gateway
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-polly
- amazon-web-services
- aws-lambda
- claude
- ffmpeg
- mcp
- or-tools
- playwright
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.