Atomic AI

Multi-tenant AI automation for personal productivity and teams: connect your apps, set plain-language rules, and let autonomous agents do the work — with humans approving anything high-stakes, instant SMS alerts online or offline, and full voice control powered by Amazon Nova Sonic for hands-free, accessible use.

Inspiration

Every individual and team lives across a dozen tools — Gmail, Slack, Jira, HubSpot, Stripe, Google Drive — and the glue between them is still a human copying context from one tab to another. For busy professionals, founders, and teams, context-switching eats hours of life every week. AI agents can act as that digital glue, but handing an autonomous agent the keys to your email, cloud infrastructure, or financial tools is terrifying: one bad tool call can send the wrong reply to a client, push a bad commit, or execute an unwanted transfer.

We wanted to build an automation layer that respects human life, time, and agency — a platform you could actually trust with your personal and professional accounts. That meant four things had to be true at once:

  1. Agents powerful enough to orchestrate complex cross-platform workflows.
  2. A collaborative human-in-the-loop approval system so nothing high-impact happens without a person explicitly saying yes.
  3. Omnichannel alerting (SMS) so you don't have to stay glued to a computer screen; whether online or offline, critical pending actions reach you anywhere.
  4. Genuine, end-to-end accessibility so that visually impaired or disabled users can navigate, review, edit, and trigger actions entirely hands-free.

Atomic AI is our answer to making AI work for humanity — safely, accessibly, and reliably.

What it does

Atomic AI is a production-grade, multi-tenant platform where individuals and teams deploy AI agents across 12 categories of apps (email, CRM, dev tools, support, calendars, cloud, and more):

  • Personal & Work Automation with Team Workspaces. Manage personal side-projects or organization workspaces with granular RBAC (Owner/Admin/Member/Viewer) that gates who can connect integrations, edit shared rules, and approve agent actions.
  • Shared & Personal Integration Vaults. Connect app accounts via OAuth and mark each connection "personal" or "workspace-shared" so agents act securely on your behalf. Every credential is encrypted at rest.
  • Plain-Language Guardrails. Write plain-English rules like "all customer emails must be saved as drafts unless an admin approves," and agents follow them rigorously.
  • Collaborative Approval Hub. When an agent proposes a high-impact action (send an email, terminate an instance, issue a refund), a BeforeToolCall interception pauses it and posts an approval request to a real-time queue. Reviewers can inspect the proposed action, edit the content, reschedule, or approve/reject it — with updates broadcast live over WebSockets.
  • Omnichannel SMS Alerting (Online & Offline). Never miss an important notification, message, or approval request. Add a phone number to your profile, and Amazon SNS sends real-time SMS alerts directly to your phone when an agent hits an approval checkpoint. You don't need to stare at a dashboard or even have an active internet connection to know when your input is needed.
  • Full Voice Control & Accessibility (Powered by Amazon Nova Sonic). Built ground-up for disabled users, including the blind and visually impaired, Atomic AI integrates Amazon Nova Sonic bidirectional speech-to-speech. Users can navigate the entire app, listen to unread emails or pending approvals read aloud, dictated word-for-word, and edit or trigger actions completely hands-free via spoken commands — including precise text adjustments (e.g., "change 'Hi there' to 'Hello PayRogen'") and executing primary UI actions ("Approve and send", "Save to draft", or "Schedule for tomorrow at 9 AM").
  • Super Admin Control Plane. System-wide user/workspace management, live agent session inspection, token/cost analytics, MCP health, and an emergency kill-switch.

Everything runs with a single docker compose up.

How we built it

  • Backend: FastAPI on Python 3.14 with an async SQLAlchemy 2.0 ORM over PostgreSQL 18, Alembic migrations, and Redis for caching, the job queue (ARQ), and real-time WebSocket coordination.
  • Agent Engine: The Strands Agents SDK driving Amazon Bedrock models, paired with the Model Context Protocol (MCP) for per-provider tools. A before_tool_call hook classifies high-impact actions and routes them to the collaborative approval queue instead of executing directly. Hard per-run caps (turns, tokens, sliding context window) keep costs bounded.
  • Voice & Speech-to-Speech Processing: Integrated Amazon Nova Sonic via the Strands BidiAgent, wrapped in a deterministic server-side Intent Router. Instead of leaving app navigation or high-stakes actions to probabilistic model tool-use, the server handles navigation routes and action triggers deterministically, allowing Nova Sonic to handle natural, expressive spoken conversation. The Next.js frontend streams 16 kHz PCM audio over WebSockets for low-latency, bidirectional audio playback.
  • Rich Interactive Editing for Voice: Integrated a headless TipTap (ProseMirror) editor that makes document and reply bodies programmatically editable, enabling voice commands to replace, insert, and rewrite text snippets accurately.
  • SMS & Offline Alerting: Utilized Amazon SNS SMS (sharing AWS credentials with Bedrock). Built a pricing/segment model estimating cost per destination country (GSM-7/UCS-2 segmenting) with a per-user monthly spend cap and best-effort sending guarantees so notification network glitches never block approval creation.
  • Frontend: Next.js 16 (App Router, TypeScript, Tailwind, shadcn UI) featuring a workspace switcher, approvals hub, rules manager, integration vault, accessibility/voice overlay, and super-admin portal.
  • Containerization: Clean multi-stage Docker Compose setup running PostgreSQL 18.6-alpine, Redis 8.10-alpine, the FastAPI backend, the ARQ background worker, and the Next.js frontend.

Challenges we ran into

  • Bridging Voice Control with Assistive Reliability. Speech-to-speech models can be probabilistic about calling tools, causing them to narrate false apologies over actions the app had already completed. We solved this by moving all navigation and action triggers to a deterministic server-side intent router, withholding those tools from the speech model, and providing grounding notes so the voice assistant never misinforms visually impaired users about execution states.
  • Handling Messy, Multi-Turn Voice Editing. Real human speech is conversational and fragmented — a blind user might say "change the text 'hi there'" … pause … then "to hello," or "approve and schedule" … "for next Monday" … "at 3 PM." We engineered per-session speech buffers that accumulate search-and-replace strings or dates/times across multi-turn exchanges to reliably modify reply bodies and schedule calendars.
  • Eliminating SMS Sandbox & Delivery Silent Failures. Cloud messaging can fail silently at the network edge due to IAM scope issues (sns:Publish) or SMS sandbox limits. We added recorded retry states, user spend logging, and verified delivery paths to ensure critical alerts reach users online or offline without fail.
  • Runaway Agent Context & Token Costs. Early inbox-processing runs re-sent prior drafts in context every turn, creating quadratic token growth. We implemented hard per-run turn/token caps, sliding context windows, durable "already-processed" item sets, and Redis poll locks.

Accomplishments that we're proud of

  • True Human-in-the-Loop Governance: Agents execute routine heavy lifting, but high-impact decisions always require human approval, leaving a transparent audit trail.
  • A Platform Built for Every Human Being: Visually impaired, blind, or physically disabled users can navigate, inspect, edit wording, schedule, approve, send, and reject workflows entirely by voice, backed by SMS alerts on their phone.
  • Never-Miss Offline Reliability: Critical notifications reach users via SMS whether they're working at their desk or completely offline on the go.
  • Deterministic Voice Orchestration: Successfully layer deterministic code logic over probabilistic speech-to-speech models to eliminate hallucinations in UI control.
  • Turnkey Production Architecture: The entire multi-tenant stack — auth, RBAC, voice SDK, encrypted vault, SMS alerts, and super-admin — boots end-to-end with a single Docker command.

What we learned

  • Accessibility Makes the Product Better for Everyone. Building hands-free voice controls and offline SMS alerts for disabled users ended up creating the ultimate productivity feature for busy founders and teams on the move.
  • Trust is the Ultimate Feature. In autonomous AI systems, approval hubs, audit trails, and strict spend caps are just as vital as the agent's intelligence.
  • Keep Probabilistic Models Out of Deterministic Loops. Voice models excel at understanding intent and natural human expression; application state, navigation, and critical database calls belong in deterministic code.

What's next for Atomic AI

  • Two-Way SMS Execution: Allow users to approve, edit, or reject agent actions directly by replying to SMS text messages with simple codes.
  • Expanded MCP Connector Ecosystem: Add deeper integrations across all 12 categories, including specialized financial, medical, and legal tool sets.
  • WhatsApp, Telegram & Push Gateways: Expand beyond SMS to offer WhatsApp Business, Telegram, and native mobile push notifications with custom quiet-hours and batching.
  • Multilingual & Proactive Voice Briefings: Enhance Amazon Nova Sonic integration to deliver proactive daily spoken briefings in multiple global languages.
  • Custom Visual Swarm Builders: Enable users to drag-and-drop custom multi-agent swarms with personalized safety rules and approval triggers.

Built With

Share this project:

Updates

Submission history