Inspiration
I start every morning the same way: weigh out the beans, grind, tamp, and wait for machine to build up the nine bars pressure to do their job. That first shot sets the tone for the rest of the day. When it lands, it's sweet and bright, and sets me for a great start of the day. When it doesn't, I'm standing at the counter with a cup that's harsh and bitter, running the checklist in my head: the beans, the grind, the dose, or me? I've been chasing espresso long enough to know the theory; the part that wears me down is the loop around it. Every new bag resets everything, and the whole research-guess-pull-taste cycle eats the morning before I've even had the coffee I was trying to make. So I built an agent to carry that loop for me (and for anyone else who wants the daily shot to just work) while I keep the one thing only I can do: taste it and decide what changes.
What it does
9Bars turns the espresso dial-in into an agentic workflow with human in the loop. Point it at some beans, a name, a roaster's page, or a photo of the bag and it reads the label, drafts a starting recipe, and shows the evidence behind it. The creates a best estimate profile for the modded Gaggia espresso machine and asks for approve of the draft profile. After the shot, it reads the telemetry, asks how it tasted, and proposes exactly one change for the next pull. Currently, It targets Gaggiuino- and GaggiMate-controlled machines, and every fact it states carries its evidence.
How I built it
- Strands Agents SDK for the agent loop, with DeepSeek as the model through strands OpenAIModel and with tool-call round-trips.
- 8 available tools over small ports (ShotSource, BrewLog, CoffeeResearcher). All SQL lives in one DuckDB adapter.
- A pure domain layer for the physics and the one-variable decision rule, grounded in the poroelastic espresso-flow paper (arXiv:2512.21528) rather than heuristics.
Brew ratio is derived: $$ r = \frac{y}{d} $$ and flow deviation is the signed distance outside the expected band: $$ \Delta f = \begin{cases} f - f_{hi} & f > f_{hi} \ f - f_{lo} & f < f_{lo} \ 0 & \text{otherwise} \end{cases} $$
Channeling is flagged, not asserted. After peak pressure we measure the drop and the flow gain, and report indicated only when both cross their thresholds: $$ \Delta p = p_{peak} - \min(p_{late}) \ge 2.0\ \text{bar}, \qquad \Delta f_{gain} = \max(f_{late}) - \min(f_{late}) \ge 1.0\ \text{g/s} $$ The one-change table is deterministic — ratio steps of 0.5, clamped at 1.0 — and each recommendation ships with a confidence: $$ c = \min\bigl(1,\ \max(0,\ 0.5 + 0.1\,\Delta)\bigr) $$
- FastAPI + SSE for the stream, a Next.js static export for the UI, and a read-only MCP server at the device boundary. The single machine write is DeviceProfileWriter, reachable only through the approval route with a server-minted token.
Challenges I ran into
- Keeping the agent from fabricating. Telemetry is a hypothesis until taste confirms it, so the tools return indicated / unconfirmed, never "channeling" as fact. The system prompt enforces the same line.
- The write path. Making the machine write safe meant removing any write tool from the agent and routing uploads through an approval token, with idempotent deploys so a double-approve can't double-write.
- DeepSeek thinking mode wasn't consistent with tool calls, so it's switched off for better latency and more accurate tool calls instead of over-thinking and burning tokens.
- The real GaggiMate binary format understanding.
Accomplishments that I am proud of
- Grounding the agent in published, verified and accurate physics.
- The taste bloom feedback loop, a bean-shaped radar for rating a shot, because the numbers and theory only take you so far and different people have different taste.
- A human-in-the-loop gate that's enforced in code, only gates the actions that need human approval.
What I learned
The interesting work on a hardware agent is the boundaries, not the model or the harness. Read-only tools, approval tokens, idempotent writes, and a "measurement vs inference" rule added more reliablity and predictibility than any prompt tuning. Another learning was that grounding in the LLM/AI agents era is a product feature.
The Strands SDK had its own gotchas:
- Thinking mode doesn't play well with tool calls. DeepSeek runs chain-of-thought by default, but strands agents doesn't round-trip reasoning_content across tool turns.
- Tool I/O is split across the trace tree. The input sits on the assistant message; the result sits on a separate tool span keyed by toolUseId. We walk the trace and join them by id to reconstruct what the agent did ( extract_spans ).
What's next for 9Bars
Probably an AgentCore deployment, and support for more machines beyond Gaggiuino and GaggiMate if I want to make this scale bigger and try to polish the user experience and test with a small group of beta testers that are also home baristas/coffee enthusiasts.

Log in or sign up for Devpost to join the conversation.