Inspiration
Support teams need automation that can investigate a payment issue without quietly authorizing its own refund. SafeOps + Strands separates an agent's proposal from permission to execute it.
What it does
A real Strands agent reads customer and payment records, identifies a duplicate payment, and proposes a refund. Its five tools route through SafeOps' generic API. The gateway checks permission, applies policy where required, assesses contextual risk, and pauses for a human when approval is needed.
In the completed historical PAY-9005 demonstration, a $750 refund reached SUPPORT_REFUND_APPROVAL with LOW risk, score 25. A human approved it. The original verification found exactly one refund and 19 audit events, with ToolRequest, ApprovalRequest, and ExternalActionRequest all EXECUTED. These are seeded demo database records, not production payment transfers.
How we built it
The submission-specific component is a real Strands Agent with thin Python tool wrappers over the existing SafeOps client. SafeOps uses FastAPI, PostgreSQL, and a Next.js operator dashboard. The historical live agent used Anthropic; Bedrock was blocked by AWS account verification. We do not claim a successful Bedrock run or AgentCore deployment.
Challenges we ran into
The model must stop at pending approval instead of retrying a refund. Keeping every tool behind the same gateway makes approval a backend decision rather than an instruction the model can override. AWS account verification prevented the Bedrock path during the demo, so the verified live run used Anthropic.
Accomplishments
The historical run exercised real model-driven investigation and an enforced human decision before changing the seeded payment record. Existing security evidence also shows a malicious support-ticket instruction blocked by contextual risk, with a denied tool request and a critical incident. This is a demonstrated scenario, not a claim of universal prompt-injection prevention.
What we learned
Useful agent autonomy and human authority can coexist when the execution boundary lives outside the model. Low risk does not cancel a policy requirement for human approval.
What's next
Future work could improve judge-accessible deployment and operational robustness. Those capabilities are not claimed as delivered by this submission.
Pre-existing work disclosure
SafeOps core, its dashboard, authorization engines, audit trail, generic external-agent API, and MCP adapter are pre-existing independent work. The submission-specific work is integrations/strands/, recorded in commit 6955009 dated September 9, 2026. The entire SafeOps platform is not a from-scratch hackathon build. Commit dates do not establish original authorship dates; eligibility remains subject to the event's new-project rules.
Evidence and limitations
Historical evidence and audit sequence
The public repository is licensed Apache-2.0. The demonstration used Strands with Anthropic, not Bedrock. Refunds update seeded demo records. The malicious email tool is a demo no-op, and the historical BLOCK case prevented its invocation. No universal exactly-once or complete prompt-injection protection is claimed. A public hosted test instance has not yet been verified.
Built With
- anthropic
- docker
- fastapi
- next.js
- postgresql
- python
- strands-agents
- typescript
Log in or sign up for Devpost to join the conversation.