Inspiration
Everyday administration takes more attention than the final decision deserves. Scheduling a car service means comparing providers, prices, ratings, and appointment times—just to make one booking.
We built DecisionPilot AI for busy individuals and families who want help with that preparation while retaining control over spending and commitments. Our principle is simple: your agent does the work; you make the call.
What it does
DecisionPilot turns an objective into a complete, auditable workflow:
“Schedule my annual car service under $180. I prefer a morning appointment and want the best balance of rating and price.”
The agent searches and evaluates candidate appointments, selects the strongest eligible option, and stops at a Human Decision Gate. The interface presents the proposed provider, appointment, price, rationale, and alternatives.
The human can approve or reject. Approval enables execution of the specific proposal, followed by receipt verification. Rejection ends the workflow without a booking. The interface displays autonomous progress, the decision state, an audit trail, and estimated human minutes saved.
The MVP uses fictional providers, illustrative prices, and simulated bookings. No garage is contacted and no payment is made. The Strands/Bedrock planning, approval controls, persistence, and verification logic are implemented and working.
How we built it
The application uses Python and the Strands Agents SDK, with Amazon Nova Lite through Amazon Bedrock.
Each planning invocation creates a fresh Strands agent with two custom tools:
search_appointmentsretrieves the synthetic appointment inventory.evaluate_appointmentscompares eligible options using rating and price.
The agent returns a structured recommendation. The controller independently validates the selected appointment against trusted inventory and the confirmed budget and morning constraint.
The public application runs on Amazon EC2, with Amazon CloudFront providing HTTPS. SQLite stores workflow state, proposals, simulated receipts, and audit events on encrypted Amazon EBS storage. AWS IAM roles provide access without embedding static credentials, and AWS Systems Manager supports host administration.
Amazon Bedrock AgentCore is an optional integration path; it is not part of the deployed demo.
The Human Decision Gate
The most important architectural choice is that the agent has no approval tool and no execution tool.
Human approval happens outside the autonomous agent loop, through a separate controller. The controller checks browser-session ownership, binds the decision to an exact proposal, enforces expiration, and rejects replayed decisions.
After approval, idempotent execution prevents duplicate simulated bookings. Verification compares the receipt against the approved provider, appointment, and price. Completion is recorded only when those checks pass.
The Human Attention Budget
DecisionPilot makes the intended benefit visible:
Estimated minutes saved = manual-effort baseline − active human interaction time
The user can adjust the manual baseline. The browser tracks active interaction time while excluding hidden tabs and extended inactivity. Savings are credited only after successful verification.
In the recorded demonstration, the interface displayed 33.8 estimated minutes saved, using a 35-minute baseline and 74 seconds of active attention. This is a transparent estimate, not an independently measured productivity result.
Challenges we faced
The hardest challenge was translating “ask the human first” into an enforceable software boundary. A prompt alone does not prevent an agent from approving its own action if that capability is available.
We addressed proposal tampering, expired approvals, replayed decisions, concurrent execution, session ownership, and false completion after receipt mismatches. We also separated durable workflow state from individual agent invocations.
Initial environment limitations made dependency installation and cloud validation difficult. GitHub Actions provided a reproducible path to validate the SDK and run an authorized live Bedrock test. We then deployed the application through a reviewed CloudFormation change set and tested approval and rejection in the public interface.
Accomplishments we’re proud of
- Built and deployed a focused car-service MVP.
- Verified real Strands/Bedrock planning through tool use to the human decision gate.
- Passed 32 automated lifecycle and HTTP tests.
- Demonstrated approval, verified completion, and rejection in the live application.
- Implemented persistent state, an audit trail, duplicate protection, and transparent attention estimates.
- Published the source under the MIT license, with deployment instructions and architecture documentation.
What we learned
Useful autonomy depends on clear boundaries and reliable recovery as much as model reasoning.
Approval must authorize a specific proposal. Execution must tolerate retries. Completion must follow verification. Time-saving claims must distinguish measured interaction time from an assumed baseline.
These lessons shaped DecisionPilot into an application where the human decision is part of the architecture.
What’s next
The next step is integrating a real service provider with authenticated user accounts, fresh quotes, provider-side idempotency, and independent receipt verification. We also plan to strengthen origin security and operational monitoring before handling real personal or payment data.
We will expand to other administrative workflows only after the car-service experience is reliable.
Try DecisionPilot AI
Built from an entrant-supplied DecisionPilot starter package with AI coding assistance. Third-party SDKs retain their respective licenses. All appointment data and bookings are synthetic.
Built With
- agents
- amazon
- amazon-cloufront
- amazon-ebs
- amazon-ec2
- aws-cloudformation
- aws-iam
- aws-system-manager
- bedrock
- nova
- python
- sdk
- strands
Log in or sign up for Devpost to join the conversation.