Juliet — Devpost fields (AWS "Agents for Humans", Professional Agents)
Numbers are in (re-verified 2026-08-14). **Where they go on Devpost:* the Results line is now the top bullet of "Accomplishments" — paste that section as-is; a one-line callout is woven into the end of "What it does" so the number shows early; and there's an eval line under "Try it out." Everything below this quote is paste-ready.*
Inspiration
The founder of a service-disabled veteran-owned small business (SDVOSB) runs an entire back office alone — chasing proposal deadlines, triaging email, hunting for new solicitations, keeping a CRM and document library current — all while trying to build a product. Most "AI assistants" are a chat box you have to babysit. We wanted the opposite: an assistant that runs in the background, does the routine work itself, and only interrupts you when there's a real decision to make.
What it does
Juliet is an autonomous executive assistant for a busy operator. She runs on a schedule, reviews open and overdue tasks, the calendar, every proposal deadline, and inbox state, and scans live federal opportunity sources (SAM.gov) for relevant solicitations. Her loop is Discover → Decide → Draft → Record: she handles everything routine and safely reversible herself, and anything that needs human judgment she raises as a structured decision (summary, context, options) in a Command Center and pushes out-of-band. You approve with one click; she carries it out with her tools and reports what she did. Outcomes and institutional knowledge accrue in a two-layer memory so she compounds over time. She only interrupts you for a real call — that restraint is the whole point, and it's measured: on a 41-item labeled triage set she scored 88% (36/41) with 0 missed escalations — she never once auto-handled something that needed a human.
How we built it
- Strands Agents SDK (TypeScript) as the agent runtime and tool loop, reasoning with Claude (Amazon Bedrock), deployed live on Amazon Bedrock AgentCore Runtime (us-east-2) — the ARM64 container built in AWS CodeBuild and pushed to ECR, no local Docker.
- Amazon EventBridge Scheduler fires the scheduled background sweep (invokes the runtime every 30 min, work hours CT); AWS Secrets Manager holds the runtime config; an IAM execution role is the runtime's identity for Bedrock.
- Supabase (Postgres + RLS + realtime) is an owned single source of truth the agent and UI share. We deliberately chose an owned Postgres over AgentCore Memory so the agent, the Command Center, and every future surface share one dataset under one access model — not state trapped inside the runtime.
- A React + Vite + Tailwind Command Center (on Cloudflare) renders the decision queue, notifications, opportunities, and call log, with realtime approve/decline.
- We re-homed an existing library of ~56 tools (tasks, calendar, Gmail, CRM, documents, proposals) onto Strands through a thin 1:1 adapter — no rewrites — and added the escalation loop, decision de-duplication, two-layer memory, and live SAM.gov Opportunities API discovery.
Challenges we ran into
- Restraint is hard. Making an agent that acts is easy; making one disciplined enough to act silently on the routine and interrupt only for genuine judgment calls took a real escalation model, de-duplication of repeat alerts, and a "fail toward surfacing" posture.
- Owning the data. We started on a managed backend that wouldn't expose database credentials — so the agent couldn't reach it. We moved everything onto a Postgres we control, which is what finally let the agent, the UI, and future surfaces share one source of truth under one access model.
- Small correctness bugs that matter. The model would occasionally guess the wrong year on a due date; we fixed it durably by injecting the current date into context at runtime.
Accomplishments that we're proud of
- A number you can re-run — 0 missed escalations. We built a re-runnable triage eval: 41 real back-office items, each hand-labeled by the operator with the correct action (handle silently / surface a decision / ignore), run against Juliet's actual decision policy on Amazon Bedrock. She scored 36/41 (88%) — and the number that matters, 0 missed escalations: she never auto-handled something that needed a human, surfacing 15/15 judgment calls including the traps (a vendor changing bank details, an auto-renewing contract, wire approvals). All 5 misses were over-surfaces — erring toward asking. Ground-truth key + scorer in the repo:
npm run eval:triage. - The escalate → approve → act → report loop working live, end to end, against a real database — she surfaces a decision, you approve it, she executes it and reports back.
- An autonomous sweep that surfaces real, correctly-prioritized decisions and degrades honestly — when a tool isn't connected she does what she can and says what she couldn't, instead of faking a result.
- ~56 existing tools re-homed onto Strands with zero rewrites.
- A branded product experience, not a proof of concept.
What we learned
- "Only surface a real decision" is a product feature, not a technical one — and it's where the trust comes from.
- A clean tool contract ports 1:1; the model-driven loop consumed our existing tools unchanged.
- Owning the database is what makes one agent, many surfaces, and one access model possible.
- Graceful degradation and honest framing beat capability breadth. An assistant that admits "I couldn't reach your calendar" is more useful — and more trusted — than one that pretends.
What's next for Juliet?
Per-teammate learning that rolls up into one org brain; a living, shared document library with private/confidential/shared/public tiers; a public-facing answering + routing capability; a team-chat surface (Slack/Teams); a proposal engine that shreds a solicitation into a compliance matrix and drafts from a canonical library; follow-through / open-loop tracking; and a customizable persona with voice and an on-screen avatar. One brain, many doors.
Built With
- ai-agents
- amazon-bedrock
- amazon-eventbridge
- amazon-web-services
- autonomous-agents
- aws-secrets-manager
- bedrock-agentcore
- claude
- cloudflare
- llm
- node.js
- postgresql
- react
- sam.gove-api
- strands-agents
- supabase
- tailwindcss
- typescript
- vite
Log in or sign up for Devpost to join the conversation.