Inspiration

I've always loved history, and I go to museums now and then. But over time I started noticing that most people around me just aren't that interested anymore, mostly kids. People walk past objects that took someone years to make, glance at the label and move on. It's a bit sad, and it feels like it's getting worse.

Around the same time I got really into Japanese culture and traditional craft, which I've been a fan of for a long time. While going down that rabbit hole I came across the karakuri chahakobi ningyo, a wooden tea-serving doll from the 1700s. You wind up a spring made of whalebone, put a cup on its tray, and it walks the tea over to you. It stops while you drink, and when you put the cup back it turns around and goes home.

What got me was that it's basically a machine doing a boring, repetitive job for a person. Which is pretty much what this hackathon is asking for, just 230 years late. ( I go this thought as I was going through the architecture of strands ).

And it made the museum problem click for me. Most museums can only show you something like this as a photo and a label. You see what it looked like but not how it actually worked, so of course people walk past it. To make it properly explorable, someone has to dig up the sources, check if they're allowed to use them, read through them, pull out measurements, figure out what moves what, and keep track of where every fact came from. That takes days per object, and it's the same slog for the next four hundred. That felt like a good job for an agent.

What it does

Echoes is a small Japanese museum that runs in your browser. Instead of looking at a photo, you get to use the objects. You can take the tea-serving doll apart and run it, play a koto with thirteen strings, and watch a bamboo shishi-odoshi fill up, tip over and clack against its stone.

Behind it is a curator agent I built with the Strands Agents SDK. It picks which licensed sources to download, actually reads the pages and photos (one of them is an illustrated manual from 1796), and works out how each object is put together. Every measurement links back to the page or image it came from, and there's a "Making Echoes" page where you can follow that trail yourself.

There's no AI in the part visitors use. It's just static files, so it loads fast and works offline.

How I built the museum

The agent only runs while I'm preparing the exhibits, never while someone's visiting. The only thing that makes it to the website is the finished, checked data. My keys, the model calls and the agent itself stay on my machine.

The curator has two tools: one to fetch a source and one to read it. I didn't want it making up citations, so before it's allowed to cite something, the code checks that the text or image really was sent to the model in that request. If it never read it, it can't cite it. Each fact is also tagged as documented, observed in a photo, inferred, or my own choice, so a guess never gets passed off as history.

The model only ever produces data, never code. My own code turns that data into the 3D objects, so the same evidence always gives the same result.

The website is Three.js. All the moving parts run off one simulation so everything stays in sync. The koto has no recordings in it at all. The sound is generated as you play using a physical model of a plucked string, and it's in tune to about a cent.

I wanted people to be able to check it's real, so there are commands for that too: delete the agent's output and the build won't run, change one of its numbers and the object changes shape, or trace any measurement back to the exact source it came from.

The part I'm proudest of: agents that build and test

The museum works, but that curator only fills in numbers for 3D code that I wrote by hand. I wanted to see if agents could build the object itself, and check each other's work honestly. So on a separate branch I built a multi-agent system with Strands, and tried it on a fourth object: a wadokei, a Japanese clock where the length of an hour changes with the seasons.

It has three agents, and the whole design is about stopping them from fooling each other.

1. The requirements agent. It reads the evidence and writes the checks the finished clock has to pass, one check at a time, sixteen in total. Once they're done they get hashed and locked. Nothing later in the pipeline is allowed to change them.

2. The builder. It gets the evidence and the locked checks, and describes the clock as structured data: every component, its size and material, where it sits, how it's joined to the others, which parts drive which, and how you interact with it. It never writes code. A separate piece of my code turns that description into the actual 3D clock.

3. The tester. This is the important one. It gets the locked checks and the finished clock, and that's it. It never sees the builder's explanation of what it made, so the builder can't talk its way into a pass. It tests the clock for real through MCP tools: one drives an actual browser, winds the handle and checks whether the right gears really turn, and another measures the geometry directly, counting parts, checking joints connect, and looking for parts that pass through each other.

When something fails, the tester sends back exactly what broke, what it expected and what it actually got. The builder fixes it, and the tester re-runs every check, not just the failed one, so a fix can't quietly break something that was already working. The builder can't touch the checks or the tests. If it thinks a check is wrong, the most it can do is file a complaint for me to look at.

It also knows when to stop: when everything passes, when the same failure keeps coming back, when fixing one thing keeps breaking another, or when the real problem is missing evidence or requirements that contradict each other.

Along the way I used a lot more of Strands than the museum needs: the multi-agent graph, native structured output, hooks to enforce attempt limits, MCP tools, checkpoints so a long run can resume instead of restarting, and Strands' own metrics and traces to see what every agent was actually doing.

Before designing the tools, I went back through everything that had gone wrong while building the museum and turned it into 81 rules, so the new agents inherited those lessons instead of rediscovering them.

Challenges

Honestly, most of this project was challenges.

Getting the multi-agent system past its first step took five tries, and every failure turned out to be the same kind of problem:

  • I asked for all sixteen checks in one go. The model kept duplicating some and forgetting others. Generating them one at a time fixed it.
  • Then two checks kept failing because the model had to guess which measurements belonged to which check. Telling it the exact measurement names fixed it.
  • The builder spent a whole attempt trying to measure a clock that didn't exist yet, because I'd given it measuring tools too early. So tools now only show up once there's something to measure.
  • The clock's frame stayed exactly the same wrong size for three rounds. It turned out my instructions said accepted parts could never change, so the builder thought it wasn't allowed to fix them.
  • And the tester made the same mistake as the first step, listing the wrong checks in its report, until I pinned those down too.

Every time, the fix was the same idea: stop asking for one big perfect answer and break it into small pieces the model can't get wrong.

Once it did run, the tester worked better than I expected. The first full clock came back with the frame 255mm tall when it should have been ~170, and the whole thing nearly double the height it should be. The tester found all of that on its own, without being told what to look for. At one point it tried to pass something it hadn't actually proved, and my deterministic checks overruled it.

Amazon Bedrock never worked on my own account. Every model said "Operation not allowed", even after I upgraded the account and opened a support ticket. I ended up testing all fifty models just to prove it was the account and not my code. A temporary workshop account let me run this system on Bedrock for a few hours, where I compared models: Qwen's 235B vision model did the building and Mistral Large 3 did the testing. The main museum runs on OpenAI with the Bedrock code written and tested offline.

It was expensive. A single round used almost three million input tokens, which way outran my budget. When I dug into it, almost none of that was the evidence. Most of it was feedback and earlier results piling up in the conversation. That points at better context management as the real fix.

And the loop didn't converge the way I wanted. After four rounds on the workshop models, not one wrong measurement had moved. So I ran the exact same pipeline with a much stronger outside model doing the building instead. The frame went from 255mm to 168, the depth from 115 to 100, and the height from 574 to 300, all on target, and the clock went from passing 4 checks to passing 15, with none failing and one still unfinished. The clock in the research build came from that run, not from the Strands agents finishing it on their own. I'd rather be upfront about that. What it does prove is that the pipeline itself was sound, and the limit was the model doing the building.

What I learned

The biggest one: the tests only catch what you tell them to look for. All sixteen checks were about sizes and gaps, and none were about how it looked. So I got a clock with perfect measurements that looked like a plastic toy until I restyled it by hand.

Then I actually played with it for a couple of minutes and found four problems nobody else had: a handle floating 8mm off its shaft, the camera getting stuck when you rotated it, missing photos on the evidence page, and dragging the clock doing the wrong thing. Sixteen checks, an independent tester and almost fifteen hundred tests had all passed right over them.

I also learned why the cheaper models couldn't fix the frame. The positions were described relative to each part's parent, so to know where something really ended up the model had to add up a chain of positions in its head, and it kept getting it wrong. My code knew the real answer. The model was guessing.

And the agent was great at the slow, careful reading and citing that I'd never want to do by hand. But deciding what something should look like, and whether it's actually right, still needed a person.

What's next

Better context management so the build loop is affordable to run. Checks for how things look, not just their size. A check that parts which should touch actually do. Giving the builder real-world positions instead of making it add them up. And running everything on Bedrock from my own account once AWS sorts it out.

Built With

Share this project:

Updates

Submission history