The problem

A port closes. A canal blocks. A supplier's factory floods. It happens somewhere in the world every single week — and when it does, a large company's control tower team handles it in minutes.

Everyone else finds out from the news.

For the small manufacturer, the importer, the family-run distributor with twelve suppliers and no logistics department, a disruption means one person losing an entire day to spreadsheets: which of my orders are stuck? which route is still open? what does the alternative cost? who do I have to call? By the time they've worked it out, the cheap alternative is gone and the expensive one is all that's left.

The tools that solve this exist. They cost six figures a year and are sold to companies with a supply-chain department. The businesses that would benefit most are precisely the ones priced out.

What we built

SupplyChain AI gives a small business the thing only large companies have: an agent that watches the supply chain around the clock and hands over a decision, already worked out.

You describe your network in plain English — "we ship electronics from Shenzhen through Singapore to a warehouse in Berlin" — and an AI agent drafts it into a digital twin: sites, lanes, distances, costs, transit times. (Or import a spreadsheet, or start from a template.)

From then on it runs by itself:

  1. It watches. Every 15 minutes, an agent scans global news, weather, and official disaster feeds (UN GDACS, USGS earthquakes, US National Weather Service) for anything touching your specific sites and lanes.
  2. It judges. A second agent grades what it finds against your network — severity, confidence, what's actually broken versus what's merely downstream. Below a high-severity threshold it logs the event and says nothing. An agent that interrupts you constantly is worse than no agent.
  3. It computes. This is the part we refused to let an AI do. Alternative routes are calculated with real graph algorithms — Dijkstra's shortest path and Yen's k-shortest paths — over your actual network. Exact added cost. Exact added days. Exact CO₂ difference.
  4. It decides. Specialised agents rank the options, price the business damage, and write an executable mitigation plan.
  5. It asks — once. One card: "Reroute via Colombo: +$1,000, +5 days, low risk — recommended" next to "Wait and monitor" and the full reasoning. Approve, reject, or snooze.

And if you've told it you trust reroutes under $1,500 and under 6 days, it just does it, logs it, and tells you afterwards.

What makes it different

The AI never invents a number. This is the core design decision, and it shapes everything.

The obvious way to build this is to describe the problem to a language model and let it answer. That produces a confident, plausible, wrong dollar figure — and nobody should move $300,000 of inventory on a guess.

So we split the system in two. Every number an operator acts on — routes, costs, transit days, blast radius, contractual penalties, carbon, days of inventory cover — is computed by deterministic, unit-tested Python. The agents do what they're genuinely good at: reading messy news, judging severity, choosing between options, and explaining the trade-off in language a human can repeat to their boss.

A side effect we didn't expect: if the AI fails completely — rate-limited, timed out, returning nonsense — the routing engine still produces the same options. The system degrades into a slightly less articulate version of itself instead of breaking. We ship that as the designed behaviour, marked "partial" with reduced confidence.

Beyond the demo

A lot of hackathon projects work on one happy path. We spent most of the build on the unglamorous half that makes it real:

  • Contracts. Paste a delivery clause from a real supplier contract and an agent extracts the lead time, grace period and liquidated damages. Every impact estimate then includes the exact penalty exposure — not "there may be penalties."
  • Inventory. "Wait and monitor" is only offered when you genuinely have the stock to survive the outage. Days of cover versus expected disruption length, checked mathematically.
  • Connectors. Pull live order and shipment data from SAP, NetSuite, Odoo, a CSV export, or any REST feed on a schedule.
  • Playbooks. Six built-in disruption responses (port closure, supplier fire, cyclone, canal blockage, customs hold, strike) that an organisation installs and edits. The agent has to follow yours, and its steps become the team's checklist.
  • Learning. Record what a decision actually cost, and future estimates recalibrate against that reality. The system's accuracy is shown openly on an Agent Ops page — including where it was wrong.
  • Stress testing. Before anything breaks: simulate a demand spike, run a war-room comparison of scenarios, or audit the whole network to find single points of failure and see how it scores against similar networks.

Design

The product's job is to not demand attention, so the interface had to earn every pixel.

We used a restrained slate-and-indigo palette, one accent colour, hairline borders, and generous space — deliberately closer to a well-made professional tool than a dashboard covered in gauges. It's fully responsive, works in light and dark mode, and installs as an app on a phone so a decision can be approved at 2 a.m. without a laptop.

For a first-time user, a guided tour spotlights every tab and control on first sign-in — 19 steps, skippable, replayable. Nobody should need a manual to understand an inbox.

The map view uses real OpenStreetMap data with street, satellite and terrain layers, so your network sits on the actual world rather than an abstract diagram.

How we built it

Frontend — Next.js 16 and TypeScript, React Flow for the network canvas, Leaflet + OpenStreetMap for the map, Tailwind with a token-driven design system so a palette change propagates everywhere at once.

Backend — Python with FastAPI. The agents are built on the Strands Agents SDK: ten specialised agents, each returning a typed Pydantic object rather than free text, orchestrated in two multi-agent graphs. In the incident graph, the routing and impact agents run in parallel and the strategist reads both results. Every agent run is traced automatically through SDK hooks, which is how the Agent Ops page exists without us writing a logging layer.

Deterministic engine — ~5,400 lines of tested Python: graph routing, blast radius, resilience auditing, demand shocks, contract penalties, inventory cover, risk scoring, and a calibration loop.

Data — Supabase (PostgreSQL) with row-level security on every table, plus organisation roles: owners and approvers can approve decisions, planners can build and simulate, viewers can only read.

AI models — provider-agnostic by design. One environment variable switches all ten agents between Amazon Bedrock, Google Gemini, any OpenAI-compatible endpoint, or a model running locally on your own machine. That mattered practically — we developed against a local model when API quotas ran out — and it matters for a small business that doesn't want its supplier data leaving its own infrastructure.

Testing — 120 automated tests (86 Python, 34 TypeScript). The routing mathematics is tested hardest, because that's the part people trust with money.

Challenges we ran into

"Affected" is not "failed." Our first version let the AI decide which sites were down, and it kept listing every downstream warehouse as broken too. The routing engine concluded half the network was gone and found no route anywhere. The fix was conceptual, not technical: split the data model into "what actually failed" and "what is affected by it", and make the event's own facts authoritative over the model's interpretation. One schema change fixed the entire pipeline — a good lesson in the bug being in the thinking, not the code.

Rerouting the wrong thing. Early on we routed around the broken link itself, which produced meaningless "+$0, +0 days" options. What an operator actually needs is the full journey re-planned end to end — from origin to destination — for every shipment whose path crossed the failure. Rewriting that detection was the difference between a demo and something useful.

Free tiers fight back. AI providers rate-limited us constantly mid-build, and one gateway rejected the structured-output format entirely. We built key rotation, automatic fallback across models, a JSON fallback when structured output is refused, and compact schemas to fit tight token limits. Annoying at the time; it's now genuinely the most resilient part of the system.

Trusting a machine with money. The hardest question wasn't technical. When should software act alone? We settled on explicit, user-set guardrails — a maximum cost, a maximum delay, a minimum confidence — plus hard rules the agent can never cross: it won't auto-approve if any shipment has no alternative route at all, or if its own analyst flagged the assessment as uncertain. Every autonomous action is audited and reported back.

What we learned

  • Decide what the AI is not allowed to do. Naming the boundary first — the model never produces a number — made every later decision obvious.
  • Typed outputs turn agents into functions. Once each agent returned a strict object, bugs stopped being "the AI said something weird" and started being ordinary, findable bugs.
  • The best agent is quiet. We spent as much effort on when not to notify as on the analysis itself.
  • Working on real data is humbling. Everything is fine until a site has no coordinates, a lane has no price, or a supplier name in the spreadsheet doesn't match the map. Most of the engineering is those cases.

Impact

There are millions of small manufacturers, importers and distributors worldwide with a genuinely fragile supply chain and no tooling at all. A single avoided stock-out — one week of a production line not stopping — is worth more than years of what this would cost to run.

More broadly, it's a template for what we think good AI products look like: let the machine do the watching and the arithmetic continuously, and give the human back the only thing that was ever actually theirs — the decision.

What's next

Live vessel tracking so shipment positions are real rather than estimated; more ERP connectors; enough real users that the resilience benchmark becomes statistically meaningful; and a marketplace where an industry association could publish playbooks that every small business in that sector gets for free.


Try it: the live demo needs no login — break a port and watch the agent work. The full app is at supplychain-ai-nine.vercel.app ([email protected] / Demo1234!), and everything is open source under MIT.

Built With

Share this project:

Updates

Submission history