Inspiration
Every operations manager we've talked to at a small manufacturer or importer finds out about a port closure, a strike or a typhoon the same way: from the news, usually at the worst time. Then a full day goes into spreadsheets — what does it hit, what do the alternatives cost, who do I call. It's repetitive, judgement-heavy work that shouldn't have to start from zero every time. We wanted an agent that does the watching and the arithmetic in the background and shows up once, with the decision already framed.
What it does
SupplyChain AI models a company's network as a digital twin — built on a canvas, described in plain English, or imported from CSV/Excel — and runs an autonomous agent over it:
- Sentinel scans news, weather and public hazard feeds (GDACS, USGS, NWS) for every site and lane, every 15 minutes, server-side.
- Analyst grades each event against your twin: severity, confidence, blast radius. Below HIGH it's just logged; nobody is pinged.
- A deterministic routing engine (Dijkstra, Yen's k-shortest) computes the real alternate lanes with exact added cost, days and carbon. The model never invents a route.
- Router ∥ Impact → Strategist run as a Strands Graph and produce a ranked set of options, revenue at risk, SLA penalties, and an executable mitigation plan.
- Contracts and inventory are checked deterministically — a paste-in penalty clause becomes an exact dollar exposure, and "wait and monitor" is only offered when stock actually covers the outage.
- The result is one card in the Decision Inbox: Reroute via Colombo (+$1,000, +5 days, low risk) · Wait · Mitigate, with rationale, sources and a confidence badge. Approve, reject or snooze — inside a policy you set, the agent approves on its own and tells you afterward.
There's also a streaming copilot that calls the same deterministic tools ("what if Suez is blocked?"), and a public no-login demo at /demo where you can fail a port and watch the whole graph run.
How we built it
- Strands Agents SDK (Python): ten agents with
@tools and Pydanticstructured_output_models; two multi-agent graphs withGraphBuilder(incident:router ∥ impact → strategist; analysis:intel → forecast ∥ scenario → strategy → report); hooks write oneagent_tracesrow per run;Agent.stream_asyncpowers the copilot;A2AServerexposes a Lane Assessor agent to other agents over the Agent-to-Agent protocol. - Amazon Bedrock:
AGENT_MODEL_PROVIDER=bedrockswaps inBedrockModelwith a cross-region inference profile — nothing else in the agents changes. The service implementsPOST /invocationsandGET /pingwith aBedrockAgentCoreAppentrypoint, so the same container deploys to Bedrock AgentCore Runtime (IAM role, least-privilege policy and an EventBridge schedule are ininfra/aws/). - Provider-agnostic by design: Gemini, Groq/xAI/any OpenAI-compatible endpoint, or a local Ollama model behind the same one environment variable — with key rotation and model fallback on rate limits.
- Deterministic core:
routing.py(blast radius, end-to-end lane detection, k-best reroutes, flow-weighting, carbon objective),inventory.py(can we wait?),contracts.py(SLA penalties),resilience.py(audit + anonymised benchmarking),demand.py(demand shocks),calibration.py(learns from recorded outcomes) — all unit-tested, no model in the loop. - Web: Next.js 16, React Flow twin + a detailed Leaflet/OpenStreetMap view (streets, satellite, terrain), Supabase (RLS, orgs & roles), SSE streaming from the agents to the browser, Decision Inbox with an autonomy policy, execution checklists and recorded outcomes, Slack one-click approvals, web push (PWA), ERP/TMS connectors (SAP/NetSuite/Odoo/CSV/REST), shipments in flight, rate cards & carrier quotes, playbooks, Agent Ops with estimate accuracy.
Challenges we ran into
- Free-tier model quotas are tiny; we built key rotation, model fallback, and a JSON-mode fallback for large nested schemas that some providers' function-calling path rejects entirely.
- "Affected" isn't "failed": the analyst initially listed downstream nodes as failed and the engine had nothing left to reroute. Splitting the two fields fixed the whole pipeline.
- Rerouting a local segment produced meaningless "+$0" candidates; we switched to computing end-to-end lanes whose baseline path actually crosses the failure.
- Keeping the model honest: route candidates are computed first and injected as data — the Router can only rank ids that already exist, never invent one.
Accomplishments that we're proud of
A complete loop — watch → assess → compute → decide → approve → remember — that runs end to end live, with every step traced and every number exact, not estimated by a language model.
What we learned
Strands' Graph plus typed outputs make multi-agent pipelines debuggable: you can read the execution order and inspect each node's object directly. The best agent products do most of their work silently and only ask a human when a real decision exists.
What's next for SupplyChain AI
Wider Bedrock AgentCore deployment across regions; a real AIS/vessel-tracking provider behind the existing tracking interface; enough production tenants that the anonymised resilience benchmark and the outcome-calibration loop become statistically meaningful; a marketplace for third-party Sentinel tools and industry playbooks.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- fastapi
- gemini
- groq
- leaflet.js
- mem0
- next.js
- openstreetmap
- openweather
- postgresql
- pydantic
- python
- railway
- react-flow
- strands-agents-sdk
- supabase
- tavily
- vercel
Log in or sign up for Devpost to join the conversation.