Manifest — Freight Brokerage's Missing Autopilot: Ten AI Agents That Catch Fraud, Negotiate Rates, and Never Skip a Check [Live Multi-Agent Swarm on AWS]
Manifest Youtube video Demo: https://youtu.be/B9Jqe0WTjXo
From Idea to Impact
Inspiration
The Agents for Humans Hackathon's Professional Agents track asks a specific question: what does a professional spend their day on that isn't actually their judgment — and can an agent take the repetitive part while leaving the judgment to them?
Freight brokerage is a textbook answer. A broker's real skill is judgment — which carrier to trust, which rate to accept, when a delay is serious enough to escalate. But the actual hours in their day go to re-typing the same shipment details into disconnected systems, manually checking whether a carrier that just emailed them is even real, eyeballing pickup-vs-delivery photos side by side, and refreshing tracking pages hoping nothing's gone wrong. None of that requires their expertise. All of it currently requires their time.
What made this project worth building wasn't just the size of that gap — it was that the Strands Agents SDK's specific primitives (Hooks, Steering, Skills, structured output) map almost exactly onto how to close it safely. A freight broker's mistakes cost real money and involve another company's trust — this isn't a domain where "the model will probably behave" is good enough. Strands gives you the tools to make the repetitive, safety-critical parts of an agent's behavior structural guarantees enforced in code, not hopes encoded in a system prompt. That's the thesis Manifest is built to prove.
The Problem
A freight broker matches companies that need to ship something with carriers who can move it. Today that means manually checking load boards, calling and emailing carriers, re-typing the same shipment details into disconnected systems, manually checking tracking pages, and — increasingly — trying to work out whether the "carrier" replying to an email is even real. None of the systems a broker touches talk to each other, and the trust-and-safety work is entirely manual.
The evidence:
- Carrier fraud & double-brokering is a $700M–$1B/year problem. TIA estimates double-brokering fraud costs carriers $700M–$1B annually; FMCSA received 8,000+ complaints in 2025, up from ~2,000 in 2021; Verisk CargoNet logged $725M in 2025 supply-chain crime losses, up 60% year-over-year — and a carrier's identity, authority, and payment details go largely unverified beyond a manual look-up.
- Manual, reactive tracking costs the industry $15B/year. ATRI: truck detention cost the industry $15B in 2023 (39.3% of deliveries detained, 135M+ driver-hours lost); individual drivers lose $11,000–$19,000/year to uncompensated detention — because delays surface late, when nobody is proactively watching.
- Repetitive cross-system re-entry burns 100–125 hours/week. A broker processing 500 loads/week loses 100–125 hours/week to duplicate data entry — retyping the same load details into every load board, TMS, and carrier portal by hand — costing a mid-size brokerage an estimated $104K–$156K annually.
- Undetected cargo damage costs $50B–$60B/year globally. Pickup and delivery photos, when taken at all, are rarely compared side by side before a claim is disputed. LTL damage claims alone run ~$2.4B/year.
These aren't edge cases. They're systematic, evidence-backed gaps that occur in freight brokerages every day — and they're solvable with a well-governed multi-agent system.
The Solution — a brokerage co-pilot that finds, vets, negotiates, watches, and reports, with a human always in the loop
Manifest is a swarm of ten specialized Strands agents, coordinated by one Orchestrator, that runs the repetitive execution of freight brokerage end to end — while every decision that touches money, commitments, or another company's trust is gated for human review, enforced in code, not left to a model's discretion.
The core insight: the hardest part of "professional agents" isn't getting a model to draft a plausible-sounding carrier email — it's making sure the thousandth repetition of that action is exactly as disciplined as the first. Strands' Hooks, Steering, and Skills exist precisely to make that a structural property of the system instead of an emergent hope. Manifest's Carrier Outreach Agent can't send an offer above its authorized ceiling no matter what the model outputs, because that check runs in real Python code before the message ever reaches the carrier — a Bedrock Guardrail is a second, independent layer on top, not the only one.
Ten agents each own exactly one job end to end — Load-Matching, Carrier Vetting & Fraud Detection, Rate Intelligence, Carrier Outreach, Voice Check-In, Document Extraction, Cargo Condition, Track-and-Trace, Customer Update, and Playbook & Lane-History — coordinated by an Orchestrator deployed live to Amazon Bedrock AgentCore Runtime, which delegates to a separately deployed Carrier Vetting Agent runtime over a genuine cross-runtime InvokeAgentRuntime call. State persists in tenant-scoped AgentCore Memory, so each broker organization gets its own isolated namespace from one shared deployment, and a load's history survives across independent invocations without anyone re-explaining it.
What It Does
| Interaction | What You Do | What Manifest Does |
|---|---|---|
| Live Dispatch | Watch a real load run end to end, then type any carrier counter-offer amount | Five agents run a real shipment — found, priced, booked — and the exact evaluate_counter_offer() logic decides accept-or-escalate against the load's real ceiling, live, from what you typed |
| Carrier Vetting & Fraud Detection | Review a flagged carrier in the Approvals queue | Real FMCSA SAFER lookups + double-brokering red-flag analysis (remit-to mismatches, unverifiable carriers) gate whether outreach can proceed autonomously — verified live on real carriers, correctly escalating for different reasons each time |
| Cargo Inspector | Upload two photos (pickup vs. delivery), or use the examples | Real Amazon Rekognition analysis of each photo independently, flagging real condition discrepancies — and correctly reporting no discrepancy on a clean pair instead of manufacturing a finding |
| Live Tracking | Click anywhere on the map to simulate a disruption | The agent reverse-geocodes the real location, recomputes a real new ETA from real distance math off that exact point, and drafts a real customer-update message reflecting it |
| Live Investigation | Speak or type a request like "open amazon with a ring camera and compare the cost" | Real Amazon Comprehend (DetectSyntax) extracts the item from your sentence; a real Amazon search opens in a new tab instantly; in parallel, a real isolated Bedrock AgentCore browser session reads live market prices via real Amazon Rekognition OCR; the agent speaks a verdict — comparing the market price to a value you declare — synthesized with real Amazon Polly neural TTS |
| Guardrails | Flip a real Bedrock Guardrail on and off, send a preset risky message | See the real dollar amount a live, deployed Bedrock Guardrail catches — or, with it off, watch the same unauthorized-commitment language get sent through unchecked |
| Audit Trail | Click any node in the live agent-swarm topology | Full history of every tool call and its reasoning — including the genuine cross-runtime call from the Orchestrator to the standalone Carrier Vetting runtime, and the AgentCore Memory recall proof |
| Approvals | Approve or reject a queued agent decision | Nothing reaches a carrier or a customer until a human acts on it — every proposal carries its full reasoning |
All of it takes real input and produces a genuinely computed result — none of these interactions replay a fixed script.
Features
| ✅ Feature | ✅ Feature |
|---|---|
| Real multi-agent orchestration (Strands SDK) | Deterministic guardrails enforced in code, not prompt hopes |
| Live Bedrock AgentCore Runtime deployment | Genuine cross-runtime agent delegation (InvokeAgentRuntime) |
| Tenant-scoped AgentCore Memory, verified live | Real, deployed Bedrock Guardrails, live-tested |
| Real Hooks (rate-limiting, call-order enforcement) | Real Steering (redirects overreach, doesn't just block) |
| Real Skills (progressive disclosure, 54% prompt cut) | Real structured output for agent-to-agent handoffs |
| Real Amazon Rekognition photo & OCR analysis | Real Amazon Comprehend syntax-based extraction |
| Real Amazon Polly voice synthesis | Voice-driven UI via the Web Speech API |
| Live, isolated cloud-browser agent sessions | Full audit trail with real reasoning, not a log file |
| Interactive, input-driven dashboard demos | Human-in-the-loop approval queue on every risky action |
How It Was Built
Agent Core — Strands Agents SDK + Amazon Bedrock AgentCore
Every agent is a Strands Agent instance: its own model choice, its own tools, its own system prompt. The Orchestrator coordinates all ten per active load and decides which paths run autonomously versus which need a human — that gate is a code-level decision based on the vetting result, never the model's discretion.
- Hooks (
BeforeToolCallEvent/AfterToolCallEvent) make repeated discipline structural:RateLimiterHookProvidercaps how many times Carrier Vetting can hit FMCSA per conversation;RequireCallFirstHookProviderenforces that Carrier Outreach always confirms load details before ever sending an offer — guaranteed by code every single time, not a prompt instruction the model might skip on repetition #47. - Steering (
SteeringHookProvider) inspects every drafted Carrier Outreach message for a class of overreach — guarantees, uncapped terms, unauthorized binding language — and redirects with specific feedback so the agent redrafts in the same turn, fixing the mistake category, not just one instance of it. - Skills (
AgentSkills/Skill.from_file) moved the Carrier Vetting Agent's fraud checklist and the Carrier Outreach Agent's counter-offer procedure out of the permanently-inlined system prompt and into progressive disclosure — a measured 54% reduction in that agent's system prompt, verified by diffing the committed change. - Structured output — typed dataclass results (
RateRecommendation,VettingResult,ExtractionResult,CargoConditionResult) let the Orchestrator feed one agent's output into the next without re-parsing free text differently each time. - Amazon Bedrock AgentCore Runtime hosts the Orchestrator and the Carrier Vetting Agent as two separately deployed live runtimes, connected by a genuine cross-runtime
InvokeAgentRuntimecall — real distributed multi-agent orchestration, confirmed withagentcore statusshowingREADYandagentcore invokerunning real tools in the cloud. - AgentCore Memory, tenant-scoped per broker organization, was verified with two separate real cloud calls where the second recalled a prior finding — a carrier's DOT number, a remit-to mismatch, a playbook note — verbatim, without re-calling any tool.
- Amazon EventBridge removes the human from the repetition loop entirely for Track-and-Trace: status checks run on a schedule, not because a broker remembered to trigger them. Only the judgment call — is this delay significant enough to escalate? — still needs the agent.
Voice-Driven Live Investigation — Comprehend + AgentCore Browser Tool + Rekognition + Polly
The dashboard's most demo-heavy feature: say or type something like "open amazon with a ring camera and compare the cost," and Manifest genuinely investigates it, live.
- Amazon Comprehend (
DetectSyntax) walks the real part-of-speech tags to pull the product out of a natural spoken sentence — chosen afterDetectEntitiesproved unreliable at tagging everyday product phrases like "ring camera" as a commercial item. - A real Amazon search opens in a new browser tab synchronously, in the same click/voice-confirm event as the trigger — browsers block
window.open()once you've crossed anawait, so this call has to happen before any network request, not after. - In parallel, a real, isolated Bedrock AgentCore Browser Tool session — a genuine AWS-managed Chromium instance — navigates to a live price listing and takes a screenshot; Amazon Rekognition's
DetectTextOCRs the real prices out of it. - The agent speaks a verdict comparing the real market price against a value you declare, synthesized with Amazon Polly neural text-to-speech and played back in the dashboard.
Governance Layer
Anything that touches money, commitments, or another company's data is enforced in code:
send_rate_offerrefuses any offer above the authorized ceiling in real Python code, before anything reaches the carrier, regardless of what the model's output says — the primary enforcement.- A live, deployed Amazon Bedrock Guardrail sits on top as a second, independent layer — verified directly against the live API to correctly block unauthorized-commitment language.
- Nothing reaches a carrier or a customer until a broker acts on it in the Approvals queue.
Frontend — Next.js Broker Dashboard
Next.js, statically exported to S3 website hosting — no CloudFront, deliberately: a static-hosted SPA has no backend secret to protect behind a CDN, so the simplest reliable path was chosen over the more "impressive-looking" one. Amazon Cognito backs an optional login; nothing on the public dashboard gates on it. Every hero interaction takes real input and computes a real result client-side or via a direct Lambda Function URL call — none of them replay a fixed script.
Infrastructure — AWS CDK
The entire stack — DynamoDB, S3, Cognito, Bedrock Guardrails, Bedrock Knowledge Bases, the AgentCore Runtime deployment, EventBridge schedules, the dashboard's S3 hosting, and the two Lambda functions backing Live Investigation and Cargo Inspector — is defined as TypeScript CDK, stack by stack, each documented at the top of its own file.
Data Sources
| Asset | Source |
|---|---|
| Carrier / fraud data | Free, public FMCSA SAFER carrier-registry API — a real external integration, not mocked |
| Cargo photos | Synthetic, Pillow-drawn test images, labeled as such everywhere they appear — not real freight photos |
| Load board / carrier portal | Two genuinely functional mock web apps built for this project (real load boards prohibit automated access under their terms of service) |
| Rate history / market data | Seeded synthetic data, structured the same way a real historical-bookings dataset would be |
| Playbook notes | Seeded synthetic broker notes, retrieved via real keyword/RAG-style search over a Bedrock Knowledge Base |
| Live investigation prices | Real, live prices read via Rekognition OCR from real shopping listings — not seeded or fabricated |
AWS Services Used
| Service | Purpose |
|---|---|
| Strands Agents SDK | Orchestration, tools, structured output, Hooks, Steering, Skills — the framework the entire agent swarm is built on |
| Amazon Bedrock AgentCore Runtime | Live deployment of the Orchestrator and the standalone Carrier Vetting Agent, with genuine cross-runtime delegation |
| Amazon Bedrock AgentCore Memory | Tenant-scoped, per-load continuity across independent invocations |
| Amazon Bedrock AgentCore Browser Tool | Real, isolated, AWS-managed Chromium sessions powering Live Investigation |
| Amazon Bedrock Guardrails | Live, deployed guardrail independently catching unauthorized-commitment language |
| Amazon Rekognition | Cargo photo condition analysis, and OCR-based live price extraction |
| Amazon Comprehend | Part-of-speech-based item extraction from natural spoken requests |
| Amazon Polly | Neural text-to-speech for the spoken price-verification verdict |
| Amazon Connect / Transcribe | Outbound voice check-in calls to carriers who go quiet by email |
| Amazon Textract | Parsing rate confirmations and BOLs, reconciled against the shipment record |
| Amazon Bedrock Knowledge Bases | Retrieval over the broker's own historical loads and playbook notes |
| AWS Lambda | Backs the Live Investigation and Cargo Inspector real-time dashboard actions |
| Amazon DynamoDB / S3 | Application data, documents, and photos |
| Amazon EventBridge | Scheduled Track-and-Trace status checks |
| Amazon Cognito | Optional broker login |
| AWS CDK | Infrastructure as code for every AWS resource in the project |
Challenges
Learning the platform's real edges (the expected unknowns)
- The deployment AWS account's Bedrock model-invocation access is blocked for Claude, Amazon Nova, and every other foundation model (
Error 002: Access to Bedrock models is not allowed for this account) — a below-default account-trust hold, not a configuration mistake. Reasoning agents run today against a temporary stand-in model, with a documented one-env-var swap back to real Claude/Nova once access clears — rather than quietly hiding the gap. - The stand-in model itself later became unavailable mid-session, most plausibly a rate/quota ceiling under heavy same-session use. Three real fixes were tried and ruled out before reaching that conclusion, and the full root-cause writeup shipped with the repo instead of a vague "known issue" note.
- Importantly, not every AWS AI service was affected — Bedrock Guardrails'
ApplyGuardrail, the AgentCore Browser Tool, Rekognition, Comprehend, and Polly all work normally in this account, verified directly rather than assumed. Several of the dashboard's most reliable live interactions are built specifically on that distinction.
A live-view feature that AWS's own bundled client couldn't reliably render
The first version of Live Investigation streamed a remote browser session's video into the dashboard via NICE DCV (bundled inside bedrock-agentcore's BrowserLiveView). On real-world, higher-latency connections, its internal license-check step threw an unhandled promise rejection and simply stopped rendering — a real, reproducible, AWS-side bug, not fixable from this repo. Rather than ship a feature that silently failed for a meaningful share of real network conditions, it was replaced entirely with a redesigned flow: a real search opens in a guaranteed-to-work new tab, and a parallel agent-side session reads real prices via Rekognition OCR instead of depending on video ever painting.
Major retailers block automated datacenter traffic — confirmed by testing, not assumed
Building the price-comparison flow meant finding a real, live product-price source the agent's browser session could reach. Amazon, eBay, and Walmart were each tested directly, and each returned its own bot-detection block page (a "Press & Hold" human-verification challenge, a generic "Sorry, something went wrong" page) to the AgentCore browser's datacenter-originated traffic. The fix: source live prices from DuckDuckGo's shopping panel, which returns real prices unblocked, while still opening the user's own real Amazon search in a real browser tab — which isn't blocked, because it isn't datacenter traffic.
A popup-blocking constraint that dictated real architecture
window.open() is silently blocked by browsers once you've crossed an await — so the "open the real search in a new tab" call has to happen synchronously, in the same call stack as the click or voice-confirm event, before any network request. That constraint shaped the entire order of operations in the Live Investigation flow.
A CDP "too many connections" error on sequential remote-browser actions
Reconnecting a PlaywrightBrowser to the same remote AgentCore session from a second Lambda invocation, before the first invocation's automation WebSocket connection had released, returned a 429 Too Many Requests — Too many connections. Diagnosed by inspecting the SDK's own type definitions for an explicit disconnect method separate from the heavier session-ending call.
Comprehend's entity detection wasn't the right tool for the job
DetectEntities doesn't reliably tag everyday product phrases ("ring camera") as a commercial item. Switching to DetectSyntax's real part-of-speech tags — finding a trigger preposition, then taking the run of noun/adjective tokens that follows — turned out to be the more reliable real extraction method.
A Bedrock Guardrail's real limitation, found by testing it live
A topic-policy DENY can recognize that a message concerns a dollar figure, but can't compare it against a dynamic, per-call ceiling — so it also flags some legitimate offers alongside the real overreach. That's exactly why the code-level ceiling check stays the primary enforcement, with the Guardrail as a genuine second layer, not the only one.
Accomplishments
- A ten-agent Strands swarm, coordinated by an Orchestrator deployed live to Amazon Bedrock AgentCore Runtime, delegating to a separately deployed Carrier Vetting Agent runtime over a real cross-runtime call — confirmed live, not just diagrammed.
- Tenant-scoped AgentCore Memory verified with real, separate cloud round trips recalling prior findings verbatim.
- A live, deployed Bedrock Guardrail verified to genuinely catch unauthorized-commitment language — and its real limitation honestly documented rather than glossed over.
- Real Hooks and Steering enforcing repeated discipline structurally, and Skills cutting a system prompt by a measured 54%.
- A voice-driven price-verification feature where every step — speech recognition, Comprehend extraction, a real isolated browser session, Rekognition OCR, and Polly speech synthesis — is a genuine, individually-tested AWS API call, rebuilt from the ground up after the first approach hit a real AWS-side reliability limit.
- Two real, honestly-documented account-level blockers (Bedrock model access, a rate-limited stand-in model) that didn't stop the project — they're worked around, explained, and left visible rather than hidden behind confident-sounding claims.
What's Next
- Swap back to real Claude/Nova reasoning the moment the account's Bedrock model-invocation hold clears — the code path already exists behind a single environment variable.
- Wire the Playbook & Lane-History Agent's Knowledge Base to real historical brokerage data, replacing the seeded synthetic notes with a broker's actual lane history.
- Let the Live Investigation flow accept a spoken declared value, not just a typed one, closing the loop on a fully voice-driven fraud check.
- Move Track-and-Trace's EventBridge schedules onto real, live shipment data, not the current demo snapshot.
- A mobile-friendly broker view, since a lot of real carrier-chasing happens away from a desk.
Built With
- amazon-connect
- amazon-rekognition
- amazon-textract
- amazon-transcribe
- amazonbedrockagentcore
- amazonbedrockguardrails
- amazonbedrockknowledgebases
- amazoncomprehend
- amazonpolly
- aws-cdk
- aws-lambda
- cognito
- dynamodb
- eventbridge
- nextjs
- python
- s3
- strandsagentssdk
- typescript
- web-speech-api
Log in or sign up for Devpost to join the conversation.