
Everyone else built an agent for a task. Armada is the agent for any human.
It does the work on its own, proves what it did instead of claiming it, and asks you exactly once — for the decision that's actually yours.

Inspiration
This hackathon is called Agents for Humans, and almost every entry answers it the same way: an agent, for a task. We wanted to answer it literally — one agent for any human.
The trouble is that "any human" is really three different people. Maya just wants her bills paid and her life admin to stop nagging her. Sara runs a Shopify store and needs a teammate who won't post to the wrong channel or refund the wrong order. Devon coordinates a community food bank and needs something closer to a small, accountable org. Show all three the same enterprise cockpit — trust ladders, departments, a treasury, governance policies — and you've told two of them "this isn't for you."
And underneath all three sits the problem nobody wants to touch: the moment an agent can act — send the email, move the money, post the announcement — the question stops being "is it smart?" and becomes "can I trust it, and will it tell me the truth?" Agents are fluent, and fluency makes a convincing "done" whether or not anything happened. A chat window has no answer for that. An agent you can't hold accountable is one you can't hand your life to.
So we didn't build another assistant. We built a governed AI workforce that onboards like a human hire — interview, probation, trust earned through performance and lost through failure, a human signature on anything risky — that reshapes to the person in front of it, does its work silently, and surfaces exactly once, for the one decision that's actually yours.

What it does
Armada is a company you staff with AI employees — and the same product a freelancer runs as a single quiet assistant. You hire them, they go on probation, they earn autonomy through real performance, and every consequential action escalates to you for a one-tap approval on your phone, your watch, in Slack, or straight out of the macOS notch.
🪞 One product, shaped to one human. A single onboarding question (focus) decides how much you ever see. Disclosure is strictly nested and enforced by tests — but governance never turns off; only its expression changes.
- Everyday (Maya): one assistant, plain language — "Ava asks before spending your money." Core screens only, no jargon, no ladders.
- Professional (Sara): a workhorse — work log, treasury, analytics, standing mandates.
- Community (Devon): the full cockpit — departments, the L0–L3 trust ladder, objectives, audit, security.
🧑💼 Hire, don't prompt. Create an AI employee with a role, a department, a manager, and a voice. A readiness assessment places them. New hires start in shadow mode: they produce proposals and touch nothing real until you graduate them.
🎚️ Autonomy is earned, then re-earned. Every employee carries a trust score that maps to a level (Probationary → Junior → Senior → Executive). Reads earn a little, successful executions earn more, failures and denials cost trust, and trust decays on a 7-day half-life. Autonomy is never a static flag.
🔏 Approve X, and only X runs. Consequential actions (external email, invoices, moving money, sharing) always require you. Approving mints a single-use, HMAC-signed execution grant bound to a SHA-256 hash of the exact parameters you saw. The agent cannot approve one thing and run another, and a grant can never be replayed.
🤖 Real work across 34 integrations. Gmail, Slack, GitHub, Stripe, Notion, HubSpot, Shopify, Linear, Klaviyo, Jira and two dozen more — each per-user and multi-tenant, full read and write. What makes full write power safe is Armada's governance, not neutered scopes.
🙋 It works silently, then raises its hand once — on your surface. A blocked or uncertain employee doesn't guess; it reaches out to the human it reports to, pauses cleanly, and resumes exactly where it left off once answered. Our five-minute film is three proofs of this one beat, each landing on a different human surface:
- Maya approves a payment with a tap on her Apple Watch.
- Sara approves an action inline in Slack.
- Devon takes an agent-initiated voice call — the agent rings, speaks in its own voice, and answers off live data.
🎯 It runs itself, and proves it. Set a standing objective. The Chief of Staff plans, delegates across the fleet, and drives it on a cadence with no human in the loop. Then an independent auditor re-checks every acceptance criterion against live data, cites concrete evidence, and reports a verified percentage the worker can never self-grade.
📞 Call your fleet. Real Amazon Nova Sonic voice calls, each employee in a distinct voice — including multi-party conference calls: add a teammate mid-call and everyone, human and AI, hears each other and speaks, one clean floor at a time.
🚫 It never fabricates. Status, numbers and outcomes may only be asserted from what a tool actually returned. An empty read is not an all-clear — the agent says "I've nothing logged on that yet; want me to check or start it?" rather than inventing "we're on track."
One gate, every surface. Web console, iPhone, Android, Apple Watch, the macOS notch app, an MCP server, an A2A agent-card endpoint, webhooks, cron and live voice all cross the same checkpoint: trust check → approval gate → credential broker → policy screen → signed audit. No surface, and no external agent, can bypass governance.


How we built it
AWS-native, end to end. The whole system holds one invariant: every action an agent takes passes through the same governance gate, no matter what triggered it.
The central design split is "the runner is the brain, the control plane is the hands." The Strands Agents SDK runner drives the Bedrock agent loop and decides what to do, but it holds no credentials and touches no database. Every tool call it emits is an authenticated call back to the control plane — the only place that enforces trust, consumes grants, screens for prompt-injection, brokers credentials, meters spend and writes the audit trail. That separation is the security model.
- Reasoning: Amazon Bedrock — Claude Haiku 4.5 for the fast path and Claude Sonnet 4.5 for heavier reasoning (cross-region "global" inference profiles) — drives every agent run, chat turn, interview, objective cycle, the independent auditor, and the conference "conductor" that decides who speaks on a call.
- Runtime: a Python Strands runner builds a
strands.Agentper hire from a control-plane tool manifest and drives the loop; a TypeScript in-process runner (also on Bedrock) stands by as a fallback so events are never dropped. - Voice: Amazon Nova Sonic via
InvokeModelWithBidirectionalStream, behind a WebSocket gateway that keeps credentials off the client and routes every tool call through the same governance path as text. - Compute & data: a single Amazon EC2 host running Docker Compose — the Next.js control plane, the Strands runner, the Nova voice gateway, PostgreSQL 16 (via Drizzle), and Caddy for automatic HTTPS. RDS-ready.
- Identity & delivery: Firebase Auth + FCM for SSO and actionable approval pushes to iPhone and Apple Watch — kept deliberately; the agent brain, reasoning, voice, data and hosting are all AWS.
- Surfaces & protocols: a Next.js web console, a Flutter app for iOS and Android, a watchOS app and a macOS notch companion, plus an MCP server and an A2A agent-card endpoint — every one behind the same gate.
A compromised model output has nothing to skip to: it can't reach a credential and it can't bypass the gate. Credentials are brokered just-in-time, per-user, per-action, and approvals mint that one-time HMAC grant bound to the exact parameters you approved.
The AWS & Bedrock stack, in full
| Capability | How Armada uses it |
|---|---|
| Amazon Bedrock — Claude Haiku 4.5 / Sonnet 4.5 | The reasoning models for every agent run, chat turn, interview, objective cycle and the independent auditor — and the conference "conductor" that routes who speaks on a call. |
| Amazon Nova Sonic | Real, low-latency voice calls and multi-party conferences over a bidirectional stream, held inside the voice gateway. |
| Strands Agents SDK | The Python runner builds a strands.Agent per hire and drives the Bedrock loop; every tool call proxies back to the control plane for enforcement. |
| Amazon EC2 + Docker Compose | One host runs the control plane, the Strands runner, the Nova voice gateway, Postgres and Caddy — reproducible from one docker compose up. |
| PostgreSQL 16 (Drizzle) | All durable state and institutional memory (pgvector / Bedrock Knowledge Bases are the natural semantic-recall upgrade). |
| Bedrock Guardrails–ready screen | Prompt-injection and policy screening at the tool gate and on every inbound event; a drop-in seam pointed at a Bedrock Guardrail. |
| Caddy | Automatic HTTPS; routes /live + /conference (WebSocket) to the voice gateway and everything else to the control plane. |
| Firebase Auth + FCM | SSO and actionable approval pushes to iPhone and Apple Watch. |
| MCP + A2A | The fleet exposed as an MCP server and an A2A agent card, so even an external agent obeys your trust rules. |
| Signed audit (Postgres) | Every governed action is logged with the acting agent's signed identity assertion and the trust level it ran at. |

How it gets better
Armada is built to improve on three loops, not to sit still:
- Every employee learns. After each run, salient facts and lessons are distilled into Postgres-backed institutional memory; before the next run they are recalled and injected, so an employee compounds context across its entire working life. And because memory can entrench a bad habit, it is explicitly ranked beneath the current order.
- Standing tracks performance. Trust is a live score: successful executions raise it, failures and denials lower it, and it decays on a 7-day half-life. Perform, and you earn more autonomy; slip, and you're pulled back — a real HR loop, not a config flag.
- The work corrects itself. Every objective cycle ends with an independent audit against live data, and the auditor's exact finding ("Issue #22 is still missing its acceptance note") is fed into the next cycle's prompt, so the worker targets precisely what remains unmet.

Over a 30-hour run this loop drove five company archetypes to a fully verified board, and the platform discovered roughly 16 of its own systematic fixes — and grew two brand-new tools — from where the audits failed.

Challenges we ran into
1. Verifying autonomous work when you cannot trust a word the agent says. This is the deepest problem in agentic AI, and where we spent the most. An agent is fluent, and fluency produces a convincing "done" whether or not anything happened — we caught one reporting it had "successfully coordinated the Ops Brief" when the document did not exist. So verification had to be state-based (judge the world through read-only APIs, never the transcript), run in a separate auditor context the worker can't influence, cite concrete evidence per criterion, and fail safe. Then verdicts had to stay stable over a moving world: we drove auditors to temperature: 0, rewrote criteria to be mechanically checkable, and separated "criterion unmet" from "audit inconclusive" — a broken audit must never erase verified reality.
2. The agent confidently lied on a live call — and we fixed honesty structurally. Asked "are we on track?", the voice agent checked its tools, found nothing logged, and then reassured the founder "we're on track" anyway — the single most dangerous failure for an agent you hand your life to. A one-off prompt patch wasn't good enough. The anti-fabrication rule now rides inside the tool result itself ("an empty result means nothing is logged — do NOT reassure"), so the guard is in context at the exact moment the model answers, across voice, chat and autonomous runs alike.
3. Porting the whole brain to AWS without touching the crown jewels. We swapped Google ADK → Strands, Gemini/Vertex → Amazon Bedrock, and Gemini Live → Nova Sonic, while keeping the provider-agnostic governance, connectors, approvals and audit untouched. Nova Sonic taught us its own rules the hard way: it responds to audio turns via voice-activity detection, not text turns, and its transcript stream can drop the tail of an answer while the audio plays on — so we drive it with real audio and reconcile transcripts against the recording.
4. Real authority over real accounts without ever letting the agent hold a credential. The security model is an architecture, not a setting. The Strands brain decides what to do but holds no tokens and touches no database; every tool call is an authenticated hop back to a control plane that is the only place trust, grants, credentials, screening, metering and audit live. The invariant that made it defensible: every surface — web, iOS, Android, watch, notch, MCP, A2A, webhook, cron or a live voice call — crosses the identical gate.
5. Making one UI honestly become three. The adaptive surface had to hide the enterprise machinery for Maya without ever weakening it. A single focus input drives every screen from one source of truth, with tests asserting disclosure stays strictly nested (everyday ⊆ professional ⊆ community) and that the core screens — work log, treasury, inbox, approvals — never disappear for anyone. Same governance engine underneath; only the vocabulary and density change.
6. Real-time multi-agent voice, and delegation that doesn't cascade. A conference call is N concurrent Nova Sonic sessions sharing one room, which demands hard floor control (exactly one speaker owns audio and transcript at a time) and sub-second routing. Underneath, a briefly-busy runner once exceeded the dispatch ack, an inline fallback dropped the child run id, and the manager re-delegated in a cascading retry storm — so we made the completion webhook the single source of truth for delegation, made delegation idempotent, and preserved the run id across the fallback.

Accomplishments that we're proud of
We proved it, we didn't just claim it. We ran a 30-hour live experiment — five company archetypes (a dev shop, a SaaS startup, an e-commerce store, an agency, a finance desk) as real standing objectives, acting on real connected accounts, metered on a real treasury, against a living world that injected fresh external events the fleet did not create.
- The board converged to a full, independently-verified all-green, every archetype at 100% verified simultaneously, across ~48 self-directed cycles each.
- The system discovered ~16 of its own systematic fixes by watching itself work, and grew two new product tools out of what the auditor found missing.
For this hackathon we took that battle-tested engine and rebuilt its entire brain on AWS — Strands + Amazon Bedrock + Nova Sonic — then pointed it at the harder question: not "can it run a company," but "can one agent genuinely serve any one human." The answer is a live app that reshapes across three people and lands the same trust beat on three different surfaces — a Watch, a Slack message, and a phone call.
And alongside it:
- Multi-party AI voice conference calls — add an employee mid-call and everyone hears each other.
- Cryptographically enforced approvals — a one-time HMAC grant bound to the exact parameters you approved.
- Four native surfaces plus two protocol surfaces (MCP + A2A) on a single governance gate.
- Real writes across 34 connectors, per-user and multi-tenant.
We believe this is the most rigorously governed autonomous AI workforce out there — and the only one where every autonomous outcome is independently verified against live data instead of self-reported, now running on AWS and adapting to any human who hires it.

What we learned
- "For humans" means adaptive, not generic. The win wasn't adding features — it was hiding them, per person, without ever weakening the engine underneath.
- A self-report is not evidence. The single most important thing we built is the auditor that refuses to believe the agent and checks the world instead.
- Honesty is a safety property, and it belongs next to the data. Putting the grounding rule inside the tool result — not just the system prompt — is what made the agent reliably tell the truth.
- Decoupling is a security feature. Splitting the brain (Strands) from the hands (the control plane) is what makes "the agent holds no credentials" true rather than aspirational.

What's next for Armada
- Managed AWS deploy: the Strands runner on AgentCore Runtime, RDS Postgres, and EventBridge for the autonomy heartbeat.
- Semantic memory: pgvector or Bedrock Knowledge Bases for recall that scales past keyword match.
- An eval CI gate: the self-verifying loop turned into a regression harness that blocks deploys on a drop in verified trajectory.
- A marketplace of pre-built employees and grantable capabilities.
- A fourth surface, pushing the adaptive model to meet even more kinds of humans exactly where they are.

Armada · It does the work. It proves the work. It asks you once.
Built on Amazon Bedrock, Amazon Nova Sonic, and the Strands Agents SDK.
Built With
- a2a
- amazon-bedrock
- amazon-ec2
- amazon-nova-sonic
- amazon-web-services
- android
- caddy
- claude-haiku-4.5
- claude-sonnet-4.5
- docker
- drizzle-orm
- firebase
- firebase-authentication
- firebase-cloud-messaging
- flutter
- ios
- macos
- mcp
- next.js
- postgresql
- python
- strands-agents-sdk
- typescript
- watchos
- websockets

Log in or sign up for Devpost to join the conversation.