Manifest — Freight Brokerage's Missing Autopilot: Ten AI Agents That Catch Fraud, Negotiate Rates, and Never Skip a Check [Live Multi-Agent Swarm on AWS]

Manifest Youtube video Demo: https://youtu.be/B9Jqe0WTjXo

From Idea to Impact

Inspiration

The Agents for Humans Hackathon's Professional Agents track asks a specific question: what does a professional spend their day on that isn't actually their judgment — and can an agent take the repetitive part while leaving the judgment to them?

Freight brokerage is a textbook answer. A broker's real skill is judgment — which carrier to trust, which rate to accept, when a delay is serious enough to escalate. But the actual hours in their day go to re-typing the same shipment details into disconnected systems, manually checking whether a carrier that just emailed them is even real, eyeballing pickup-vs-delivery photos side by side, and refreshing tracking pages hoping nothing's gone wrong. None of that requires their expertise. All of it currently requires their time.

What made this project worth building wasn't just the size of that gap — it was that the Strands Agents SDK's specific primitives (Hooks, Steering, Skills, structured output) map almost exactly onto how to close it safely. A freight broker's mistakes cost real money and involve another company's trust — this isn't a domain where "the model will probably behave" is good enough. Strands gives you the tools to make the repetitive, safety-critical parts of an agent's behavior structural guarantees enforced in code, not hopes encoded in a system prompt. That's the thesis Manifest is built to prove.

The Problem

A freight broker matches companies that need to ship something with carriers who can move it. Today that means manually checking load boards, calling and emailing carriers, re-typing the same shipment details into disconnected systems, manually checking tracking pages, and — increasingly — trying to work out whether the "carrier" replying to an email is even real. None of the systems a broker touches talk to each other, and the trust-and-safety work is entirely manual.

The evidence:

  • Carrier fraud & double-brokering is a $700M–$1B/year problem. TIA estimates double-brokering fraud costs carriers $700M–$1B annually; FMCSA received 8,000+ complaints in 2025, up from ~2,000 in 2021; Verisk CargoNet logged $725M in 2025 supply-chain crime losses, up 60% year-over-year — and a carrier's identity, authority, and payment details go largely unverified beyond a manual look-up.
  • Manual, reactive tracking costs the industry $15B/year. ATRI: truck detention cost the industry $15B in 2023 (39.3% of deliveries detained, 135M+ driver-hours lost); individual drivers lose $11,000–$19,000/year to uncompensated detention — because delays surface late, when nobody is proactively watching.
  • Repetitive cross-system re-entry burns 100–125 hours/week. A broker processing 500 loads/week loses 100–125 hours/week to duplicate data entry — retyping the same load details into every load board, TMS, and carrier portal by hand — costing a mid-size brokerage an estimated $104K–$156K annually.
  • Undetected cargo damage costs $50B–$60B/year globally. Pickup and delivery photos, when taken at all, are rarely compared side by side before a claim is disputed. LTL damage claims alone run ~$2.4B/year.

These aren't edge cases. They're systematic, evidence-backed gaps that occur in freight brokerages every day — and they're solvable with a well-governed multi-agent system.

The Solution — a brokerage co-pilot that finds, vets, negotiates, watches, and reports, with a human always in the loop

Manifest is a swarm of ten specialized Strands agents, coordinated by one Orchestrator, that runs the repetitive execution of freight brokerage end to end — while every decision that touches money, commitments, or another company's trust is gated for human review, enforced in code, not left to a model's discretion.

The core insight: the hardest part of "professional agents" isn't getting a model to draft a plausible-sounding carrier email — it's making sure the thousandth repetition of that action is exactly as disciplined as the first. Strands' Hooks, Steering, and Skills exist precisely to make that a structural property of the system instead of an emergent hope. Manifest's Carrier Outreach Agent can't send an offer above its authorized ceiling no matter what the model outputs, because that check runs in real Python code before the message ever reaches the carrier — a Bedrock Guardrail is a second, independent layer on top, not the only one.

Ten agents each own exactly one job end to end — Load-Matching, Carrier Vetting & Fraud Detection, Rate Intelligence, Carrier Outreach, Voice Check-In, Document Extraction, Cargo Condition, Track-and-Trace, Customer Update, and Playbook & Lane-History — coordinated by an Orchestrator deployed live to Amazon Bedrock AgentCore Runtime, which delegates to a separately deployed Carrier Vetting Agent runtime over a genuine cross-runtime InvokeAgentRuntime call. State persists in tenant-scoped AgentCore Memory, so each broker organization gets its own isolated namespace from one shared deployment, and a load's history survives across independent invocations without anyone re-explaining it.


What It Does

Interaction What You Do What Manifest Does
Live Dispatch Watch a real load run end to end, then type any carrier counter-offer amount Five agents run a real shipment — found, priced, booked — and the exact evaluate_counter_offer() logic decides accept-or-escalate against the load's real ceiling, live, from what you typed
Carrier Vetting & Fraud Detection Review a flagged carrier in the Approvals queue Real FMCSA SAFER lookups + double-brokering red-flag analysis (remit-to mismatches, unverifiable carriers) gate whether outreach can proceed autonomously — verified live on real carriers, correctly escalating for different reasons each time
Cargo Inspector Upload two photos (pickup vs. delivery), or use the examples Real Amazon Rekognition analysis of each photo independently, flagging real condition discrepancies — and correctly reporting no discrepancy on a clean pair instead of manufacturing a finding
Live Tracking Click anywhere on the map to simulate a disruption The agent reverse-geocodes the real location, recomputes a real new ETA from real distance math off that exact point, and drafts a real customer-update message reflecting it
Live Investigation Speak or type a request like "open amazon with a ring camera and compare the cost" Real Amazon Comprehend (DetectSyntax) extracts the item from your sentence; a real Amazon search opens in a new tab instantly; in parallel, a real isolated Bedrock AgentCore browser session reads live market prices via real Amazon Rekognition OCR; the agent speaks a verdict — comparing the market price to a value you declare — synthesized with real Amazon Polly neural TTS
Guardrails Flip a real Bedrock Guardrail on and off, send a preset risky message See the real dollar amount a live, deployed Bedrock Guardrail catches — or, with it off, watch the same unauthorized-commitment language get sent through unchecked
Audit Trail Click any node in the live agent-swarm topology Full history of every tool call and its reasoning — including the genuine cross-runtime call from the Orchestrator to the standalone Carrier Vetting runtime, and the AgentCore Memory recall proof
Approvals Approve or reject a queued agent decision Nothing reaches a carrier or a customer until a human acts on it — every proposal carries its full reasoning

All of it takes real input and produces a genuinely computed result — none of these interactions replay a fixed script.


Features

✅ Feature ✅ Feature
Real multi-agent orchestration (Strands SDK) Deterministic guardrails enforced in code, not prompt hopes
Live Bedrock AgentCore Runtime deployment Genuine cross-runtime agent delegation (InvokeAgentRuntime)
Tenant-scoped AgentCore Memory, verified live Real, deployed Bedrock Guardrails, live-tested
Real Hooks (rate-limiting, call-order enforcement) Real Steering (redirects overreach, doesn't just block)
Real Skills (progressive disclosure, 54% prompt cut) Real structured output for agent-to-agent handoffs
Real Amazon Rekognition photo & OCR analysis Real Amazon Comprehend syntax-based extraction
Real Amazon Polly voice synthesis Voice-driven UI via the Web Speech API
Live, isolated cloud-browser agent sessions Full audit trail with real reasoning, not a log file
Interactive, input-driven dashboard demos Human-in-the-loop approval queue on every risky action

How It Was Built

Agent Core — Strands Agents SDK + Amazon Bedrock AgentCore

Every agent is a Strands Agent instance: its own model choice, its own tools, its own system prompt. The Orchestrator coordinates all ten per active load and decides which paths run autonomously versus which need a human — that gate is a code-level decision based on the vetting result, never the model's discretion.

  • Hooks (BeforeToolCallEvent/AfterToolCallEvent) make repeated discipline structural: RateLimiterHookProvider caps how many times Carrier Vetting can hit FMCSA per conversation; RequireCallFirstHookProvider enforces that Carrier Outreach always confirms load details before ever sending an offer — guaranteed by code every single time, not a prompt instruction the model might skip on repetition #47.
  • Steering (SteeringHookProvider) inspects every drafted Carrier Outreach message for a class of overreach — guarantees, uncapped terms, unauthorized binding language — and redirects with specific feedback so the agent redrafts in the same turn, fixing the mistake category, not just one instance of it.
  • Skills (AgentSkills/Skill.from_file) moved the Carrier Vetting Agent's fraud checklist and the Carrier Outreach Agent's counter-offer procedure out of the permanently-inlined system prompt and into progressive disclosure — a measured 54% reduction in that agent's system prompt, verified by diffing the committed change.
  • Structured output — typed dataclass results (RateRecommendation, VettingResult, ExtractionResult, CargoConditionResult) let the Orchestrator feed one agent's output into the next without re-parsing free text differently each time.
  • Amazon Bedrock AgentCore Runtime hosts the Orchestrator and the Carrier Vetting Agent as two separately deployed live runtimes, connected by a genuine cross-runtime InvokeAgentRuntime call — real distributed multi-agent orchestration, confirmed with agentcore status showing READY and agentcore invoke running real tools in the cloud.
  • AgentCore Memory, tenant-scoped per broker organization, was verified with two separate real cloud calls where the second recalled a prior finding — a carrier's DOT number, a remit-to mismatch, a playbook note — verbatim, without re-calling any tool.
  • Amazon EventBridge removes the human from the repetition loop entirely for Track-and-Trace: status checks run on a schedule, not because a broker remembered to trigger them. Only the judgment call — is this delay significant enough to escalate? — still needs the agent.

Voice-Driven Live Investigation — Comprehend + AgentCore Browser Tool + Rekognition + Polly

The dashboard's most demo-heavy feature: say or type something like "open amazon with a ring camera and compare the cost," and Manifest genuinely investigates it, live.

  • Amazon Comprehend (DetectSyntax) walks the real part-of-speech tags to pull the product out of a natural spoken sentence — chosen after DetectEntities proved unreliable at tagging everyday product phrases like "ring camera" as a commercial item.
  • A real Amazon search opens in a new browser tab synchronously, in the same click/voice-confirm event as the trigger — browsers block window.open() once you've crossed an await, so this call has to happen before any network request, not after.
  • In parallel, a real, isolated Bedrock AgentCore Browser Tool session — a genuine AWS-managed Chromium instance — navigates to a live price listing and takes a screenshot; Amazon Rekognition's DetectText OCRs the real prices out of it.
  • The agent speaks a verdict comparing the real market price against a value you declare, synthesized with Amazon Polly neural text-to-speech and played back in the dashboard.

Governance Layer

Anything that touches money, commitments, or another company's data is enforced in code:

  • send_rate_offer refuses any offer above the authorized ceiling in real Python code, before anything reaches the carrier, regardless of what the model's output says — the primary enforcement.
  • A live, deployed Amazon Bedrock Guardrail sits on top as a second, independent layer — verified directly against the live API to correctly block unauthorized-commitment language.
  • Nothing reaches a carrier or a customer until a broker acts on it in the Approvals queue.

Frontend — Next.js Broker Dashboard

Next.js, statically exported to S3 website hosting — no CloudFront, deliberately: a static-hosted SPA has no backend secret to protect behind a CDN, so the simplest reliable path was chosen over the more "impressive-looking" one. Amazon Cognito backs an optional login; nothing on the public dashboard gates on it. Every hero interaction takes real input and computes a real result client-side or via a direct Lambda Function URL call — none of them replay a fixed script.

Infrastructure — AWS CDK

The entire stack — DynamoDB, S3, Cognito, Bedrock Guardrails, Bedrock Knowledge Bases, the AgentCore Runtime deployment, EventBridge schedules, the dashboard's S3 hosting, and the two Lambda functions backing Live Investigation and Cargo Inspector — is defined as TypeScript CDK, stack by stack, each documented at the top of its own file.


Data Sources

Asset Source
Carrier / fraud data Free, public FMCSA SAFER carrier-registry API — a real external integration, not mocked
Cargo photos Synthetic, Pillow-drawn test images, labeled as such everywhere they appear — not real freight photos
Load board / carrier portal Two genuinely functional mock web apps built for this project (real load boards prohibit automated access under their terms of service)
Rate history / market data Seeded synthetic data, structured the same way a real historical-bookings dataset would be
Playbook notes Seeded synthetic broker notes, retrieved via real keyword/RAG-style search over a Bedrock Knowledge Base
Live investigation prices Real, live prices read via Rekognition OCR from real shopping listings — not seeded or fabricated

AWS Services Used

Service Purpose
Strands Agents SDK Orchestration, tools, structured output, Hooks, Steering, Skills — the framework the entire agent swarm is built on
Amazon Bedrock AgentCore Runtime Live deployment of the Orchestrator and the standalone Carrier Vetting Agent, with genuine cross-runtime delegation
Amazon Bedrock AgentCore Memory Tenant-scoped, per-load continuity across independent invocations
Amazon Bedrock AgentCore Browser Tool Real, isolated, AWS-managed Chromium sessions powering Live Investigation
Amazon Bedrock Guardrails Live, deployed guardrail independently catching unauthorized-commitment language
Amazon Rekognition Cargo photo condition analysis, and OCR-based live price extraction
Amazon Comprehend Part-of-speech-based item extraction from natural spoken requests
Amazon Polly Neural text-to-speech for the spoken price-verification verdict
Amazon Connect / Transcribe Outbound voice check-in calls to carriers who go quiet by email
Amazon Textract Parsing rate confirmations and BOLs, reconciled against the shipment record
Amazon Bedrock Knowledge Bases Retrieval over the broker's own historical loads and playbook notes
AWS Lambda Backs the Live Investigation and Cargo Inspector real-time dashboard actions
Amazon DynamoDB / S3 Application data, documents, and photos
Amazon EventBridge Scheduled Track-and-Trace status checks
Amazon Cognito Optional broker login
AWS CDK Infrastructure as code for every AWS resource in the project

Challenges

Learning the platform's real edges (the expected unknowns)

  • The deployment AWS account's Bedrock model-invocation access is blocked for Claude, Amazon Nova, and every other foundation model (Error 002: Access to Bedrock models is not allowed for this account) — a below-default account-trust hold, not a configuration mistake. Reasoning agents run today against a temporary stand-in model, with a documented one-env-var swap back to real Claude/Nova once access clears — rather than quietly hiding the gap.
  • The stand-in model itself later became unavailable mid-session, most plausibly a rate/quota ceiling under heavy same-session use. Three real fixes were tried and ruled out before reaching that conclusion, and the full root-cause writeup shipped with the repo instead of a vague "known issue" note.
  • Importantly, not every AWS AI service was affected — Bedrock Guardrails' ApplyGuardrail, the AgentCore Browser Tool, Rekognition, Comprehend, and Polly all work normally in this account, verified directly rather than assumed. Several of the dashboard's most reliable live interactions are built specifically on that distinction.

A live-view feature that AWS's own bundled client couldn't reliably render

The first version of Live Investigation streamed a remote browser session's video into the dashboard via NICE DCV (bundled inside bedrock-agentcore's BrowserLiveView). On real-world, higher-latency connections, its internal license-check step threw an unhandled promise rejection and simply stopped rendering — a real, reproducible, AWS-side bug, not fixable from this repo. Rather than ship a feature that silently failed for a meaningful share of real network conditions, it was replaced entirely with a redesigned flow: a real search opens in a guaranteed-to-work new tab, and a parallel agent-side session reads real prices via Rekognition OCR instead of depending on video ever painting.

Major retailers block automated datacenter traffic — confirmed by testing, not assumed

Building the price-comparison flow meant finding a real, live product-price source the agent's browser session could reach. Amazon, eBay, and Walmart were each tested directly, and each returned its own bot-detection block page (a "Press & Hold" human-verification challenge, a generic "Sorry, something went wrong" page) to the AgentCore browser's datacenter-originated traffic. The fix: source live prices from DuckDuckGo's shopping panel, which returns real prices unblocked, while still opening the user's own real Amazon search in a real browser tab — which isn't blocked, because it isn't datacenter traffic.

A popup-blocking constraint that dictated real architecture

window.open() is silently blocked by browsers once you've crossed an await — so the "open the real search in a new tab" call has to happen synchronously, in the same call stack as the click or voice-confirm event, before any network request. That constraint shaped the entire order of operations in the Live Investigation flow.

A CDP "too many connections" error on sequential remote-browser actions

Reconnecting a PlaywrightBrowser to the same remote AgentCore session from a second Lambda invocation, before the first invocation's automation WebSocket connection had released, returned a 429 Too Many Requests — Too many connections. Diagnosed by inspecting the SDK's own type definitions for an explicit disconnect method separate from the heavier session-ending call.

Comprehend's entity detection wasn't the right tool for the job

DetectEntities doesn't reliably tag everyday product phrases ("ring camera") as a commercial item. Switching to DetectSyntax's real part-of-speech tags — finding a trigger preposition, then taking the run of noun/adjective tokens that follows — turned out to be the more reliable real extraction method.

A Bedrock Guardrail's real limitation, found by testing it live

A topic-policy DENY can recognize that a message concerns a dollar figure, but can't compare it against a dynamic, per-call ceiling — so it also flags some legitimate offers alongside the real overreach. That's exactly why the code-level ceiling check stays the primary enforcement, with the Guardrail as a genuine second layer, not the only one.


Accomplishments

  • A ten-agent Strands swarm, coordinated by an Orchestrator deployed live to Amazon Bedrock AgentCore Runtime, delegating to a separately deployed Carrier Vetting Agent runtime over a real cross-runtime call — confirmed live, not just diagrammed.
  • Tenant-scoped AgentCore Memory verified with real, separate cloud round trips recalling prior findings verbatim.
  • A live, deployed Bedrock Guardrail verified to genuinely catch unauthorized-commitment language — and its real limitation honestly documented rather than glossed over.
  • Real Hooks and Steering enforcing repeated discipline structurally, and Skills cutting a system prompt by a measured 54%.
  • A voice-driven price-verification feature where every step — speech recognition, Comprehend extraction, a real isolated browser session, Rekognition OCR, and Polly speech synthesis — is a genuine, individually-tested AWS API call, rebuilt from the ground up after the first approach hit a real AWS-side reliability limit.
  • Two real, honestly-documented account-level blockers (Bedrock model access, a rate-limited stand-in model) that didn't stop the project — they're worked around, explained, and left visible rather than hidden behind confident-sounding claims.

What's Next

  • Swap back to real Claude/Nova reasoning the moment the account's Bedrock model-invocation hold clears — the code path already exists behind a single environment variable.
  • Wire the Playbook & Lane-History Agent's Knowledge Base to real historical brokerage data, replacing the seeded synthetic notes with a broker's actual lane history.
  • Let the Live Investigation flow accept a spoken declared value, not just a typed one, closing the loop on a fully voice-driven fraud check.
  • Move Track-and-Trace's EventBridge schedules onto real, live shipment data, not the current demo snapshot.
  • A mobile-friendly broker view, since a lot of real carrier-chasing happens away from a desk.

Built With

  • amazon-connect
  • amazon-rekognition
  • amazon-textract
  • amazon-transcribe
  • amazonbedrockagentcore
  • amazonbedrockguardrails
  • amazonbedrockknowledgebases
  • amazoncomprehend
  • amazonpolly
  • aws-cdk
  • aws-lambda
  • cognito
  • dynamodb
  • eventbridge
  • nextjs
  • python
  • s3
  • strandsagentssdk
  • typescript
  • web-speech-api
Share this project:

Updates

Submission history