Inspiration
A technician can replace the device perfectly and still leave the customer unable to work.
That is the daily reality for field-service coordinators. A ticket arrives — "Urgent—replace the printer. Send someone today." — and looks routine. It is not. Before anyone rolls a truck, someone has to reconcile the contract SLA against the requested urgency, confirm site access, find out who actually uses the device, surface the replacement's registration and workflow-restoration dependencies, and define what "done" provably means. That reconciliation is invisible work, and it is exactly where promises get broken.
FieldBridge was built from generalized field-service experience to answer one question: can an agent do that hidden investigation and hand a human one evidence-linked decision — without ever being trusted to act? Every organization, ticket, and identifier in the project is synthetic, so the idea could be explored without touching real customer data.
What it does
FieldBridge investigates bounded synthetic field-service tickets with evidence-specific tools, deterministic policy, and human review. It returns one decision card with the SLA deadline, contracted team, internal route, missing evidence, site-access constraints, fulfillment plan, restoration steps, closure proof, and blocked actions.
The flagship scenario starts with "Urgent—replace the printer. Send someone today." The system discovers that the actual user and serial evidence are missing, a site chaperone is required, and replacement requires network registration plus managed-workflow restoration. A second fax scenario follows a different routing path without inventing a printer serial.
The agent may inspect, route, recommend, prepare, and record a review. It cannot dispatch, order parts, contact a customer, modify production, or close work. Information requests are visibly drafts. Unknown evidence remains unknown.
How we built it
- Python 3.12, FastAPI, and typed immutable records
- Strands Agents with narrow synthetic evidence tools
- Amazon Nova Lite through Amazon Bedrock
- Amazon Bedrock AgentCore Runtime
- API Gateway/Lambda façade, WAF, DynamoDB review ledger, Secrets Manager HMAC quotas, and CloudWatch observability
- Deterministic policy validation beneath model output
- Pytest, Ruff, coverage, pip-audit, release validation, and 15 synthetic product evaluations
Responsible-agent design runs through the whole stack. The browser accepts only approved synthetic scenario IDs, so the model never sees a free-form prompt or real ticket, and it cannot change policy or reach operational tools. If model execution fails, the result is labeled DEGRADED_REVIEW_REQUIRED and the application does not fabricate a live trace. Human approval records review of a synthetic draft only.
Challenges we ran into
- Keeping unknowns unknown. Given a replacement ticket with no serial on record, the model's instinct is to be helpful by inventing a plausible one. Tool and prompt design had to make "unknown" a first-class, visible answer — the fax scenario must route differently without ever manufacturing printer evidence.
- Making deterministic code authoritative. "Send someone today" is urgency language, not a schedule. The contract SLA — not the model — computes the deadline, routing, owners, and blocked actions, which meant building a typed policy layer the model operates beneath rather than alongside.
- Failing closed without theater. When Bedrock or AgentCore is unavailable, the easiest path is a cached "successful" run. FieldBridge instead returns a deterministic packet labeled
DEGRADED_REVIEW_REQUIREDand disables approval rather than fabricating a tool trace. - Putting an agent on the public internet. The live URL needed strict schemas, approved scenario IDs only, WAF, request-size limits, per-caller HMAC quotas, idempotency, TTL metadata, and generic errors — the façade layer took real iteration to get right.
- Verifying instead of asserting. A single passing run proves little. The AgentCore deployment was verified with live Nova Lite runs and X-Ray traces, and the evaluation suite had to pass three consecutive attempts before it was counted as evidence.
Accomplishments that we're proud of
- 15/15 synthetic product evaluations passing across three consecutive runs — covering authoritative SLA/routing/safety, unknown preservation, replacement dependencies, input rejection, idempotency, and review-state behavior.
- A live, public deployment on Amazon Bedrock AgentCore running real Nova Lite inference — not a mocked demo — with captured trace evidence.
- A hard, enforced action boundary: allowed is inspect → route → recommend → prepare → record review; blocked is dispatch, ordering, customer contact, production changes, and closure.
- An append-only audit ledger where approval records review of a draft — a review button that is never secretly an action button.
- Engineering discipline throughout: pytest, Ruff, coverage, pip-audit, release validation, a Playwright test of the critical browser flow, and passing CI on the public repo.
What we learned
- The model interprets; code keeps the promise. The agent was most reliable when it chose which evidence to inspect and explained the result, while deterministic policy owned every number, route, and permission.
- Approval is not action. Framing "Approve draft" as a recorded human review of a packet — not a dispatch trigger — shaped the entire interface and audit design.
- The most valuable interruption is the right one. An agent earns trust when it interrupts only for a fact or authority it cannot safely resolve, and stays quiet otherwise.
- Honest degradation beats fabricated success. Showing a clearly labeled degraded result built more confidence than any simulated run would have.
- Evidence discipline changes claims. Requiring repeatable runs and captured traces before calling something "verified" kept the submission honest — and made the verified claims stronger.
What's next for FieldBridge
- Broaden scenario coverage beyond the printer and fax paths — network, access, and multi-site tickets that stress different evidence combinations.
- Expand the synthetic evaluation suite in step, so every new scenario type arrives with unknown-preservation and boundary tests rather than afterthought checks.
- Model richer closure proof, so "done" can be evidenced end-to-end from workflow verification through the reviewed customer update.
- Deepen the human-review experience: diffing corrections against the agent's draft and making the audit timeline easier to scan.
- Keep the synthetic-only boundary intact until any real integration is explicitly authorized — growth without overclaiming.
Built With
- agents
- amazon
- amazon-web-services
- strands

Log in or sign up for Devpost to join the conversation.