Inspiration
Most digital workflows today suffer from heavy cognitive friction. As students and developers, we kept running into the same frustrating paradox: while modern large language models can answer almost any hypothetical question, they still leave humans acting as manual middleware. Users are stuck copying values across tabs, validating constraints by hand, and babysitting fragmented interfaces.
We built our project around a straightforward premise: software should adapt to human objectives, not force humans to adapt to software interfaces. We set out to move beyond passive chatbots and build an autonomous system that takes high-level, unstructured intent, formulates a verifiable multi-step plan, and executes it with deterministic reliability.
What We Learned
Building an end-to-end agentic system shifted our entire perspective on applied AI. We realized that raw model intelligence matters far less than the orchestration architecture wrapped around it:
- Reliability Demands Constraints: Pure prompt engineering cannot guarantee task completion; deterministic schemas and strictly typed tool outputs are mandatory for enterprise reliability.
- The Psychology of Trust: Autonomous agents cannot be opaque black boxes. Exposing intermediate reasoning states, tool dispatches, and verification gates directly increases user adoption.
- Orchestration Lifecycle: Working hands-on with the Strands Agents SDK showed us how to manage session state across complex recursive loops while maintaining predictable context window budgets.
How We Built It
Our stack combines an interactive, low-latency frontend with an agentic backend orchestrated via the Strands Agents SDK:
- Core Orchestration: We utilized the Strands Agents SDK to drive the perception-reasoning-action loop. The agent parses raw natural language, constructs an execution plan, and calls targeted domain tools dynamically.
- Deterministic Tool Invocation: All tool signatures enforce strict parameter schemas. When the model invokes a capability, the parameters are validated prior to execution, preventing runtime failures.
- State & Telemetry Streaming: Execution state is persisted across conversation turns. The backend streams real-time intermediate agent thoughts, function execution traces, and output payloads straight to the user client.
- Human-in-the-Loop Safeguards: For any action exceeding a risk threshold, the engine pauses the execution cycle and yields control back to the user for explicit confirmation before resuming.
Mathematically, we framed the agent's multi-step decision process over discrete execution steps \(t \in {1, \dots, T}\). At each stage, the agent selects an action \(a_t \in \mathcal{A}\) conditioned on user intent \(x\) and historical observations \(o_{<t}\):
$$ a_t = \arg\max_{a \in \mathcal{A}} P(a \mid x, o_1, a_1, \dots, o_{t-1}) $$
Each executed action returns an environment state \(o_t\), allowing the agent to continuously self-correct until the termination criteria \(\Omega\) is met.
Challenges We Faced
- Compound Error in Multi-Step Chains: In multi-step sequences, errors propagate multiplicatively. If each independent step has a success probability \(p_i\), total chain success across \(n\) actions decays exponentially:
$$ P_{\text{success}} = \prod_{i=1}^{n} p_i $$
Without intervention, an agent attempting a 5-step task with \(p_i = 0.85\) succeeds only \(\approx 44\%\) of the time. We solved this by implementing strict schema validation gates and automated fallback retries using the Strands SDK, ensuring \(p_i \to 0.99\) at each validation boundary.
- Latency vs. Transparency: Multi-step autonomous planning introduces response latency. To eliminate user drop-off, we built an asynchronous event pipeline that streams real-time agent telemetry, letting users watch the agent think, verify, and execute.
- Balancing Autonomy with Control: Determining the boundary between independent action and human oversight required designing granular safety policies: low-risk data aggregation runs autonomously, while state-mutating operations trigger confirmation requests.
Built With
- api
- human-in-the-loop
- json
- llm
- next.js
- node.js
- prompt-engineering
- python
- react
- rest-api
- strands-agent
- tailwind-css
- typescript

Log in or sign up for Devpost to join the conversation.