Tribe.AI: build a team of AI agents you can actually watch work.
Inspiration
Most "AI assistant" demos are a chat box. You type, text comes back, and you have no idea what happened in between: which tool ran, which model was used, whether it actually opened the page it claims to have read. We wanted assistants you could watch work, the way you would watch colleagues across the office.
So we built Tribe: a pixel-art isometric office where each agent is a character at a desk. An agent stands and faces you while it listens and thinks. When it needs a computer (a shell, a file, a browser) it walks to its chair and sits down, the monitor lights up, and the real web page it is looking at is painted onto the screen. When it checks your mail it walks to the letterbox. Nothing in the room is a timer or a fake animation. Every movement is driven by a real event from the agent runtime.
What it does
- Build your tribe in the Foundry. Pick a role template (researcher, writer, scheduler, analyst, operator), give the agent a name and a look, and buy it a computer. The rig is not decoration: a Salvaged Terminal runs Haiku and can only talk, an Office Workstation runs Sonnet and unlocks a sandboxed VM and a headless browser, an Overclocked Rig runs Opus with extended thinking. What an agent owns decides what it can do.
- Access passes. Gmail and Google Calendar are linked once per account and granted per agent, so the boss can read your inbox while the writer never sees it. There is deliberately no send tool: the agent drafts, the human sends.
- A tribe of up to four. The boss hands work to colleagues using structured handoff envelopes (brief, context, deliverable, constraints) and fans independent subtasks out in parallel. Every busy colleague animates at its own desk.
- Real work products. A cloud sandbox with Python and ~200 packages builds PDFs, decks, spreadsheets and charts, which land in a per-user filing cabinet on S3 and come back as download links.
- Autonomy. Routines run on a schedule or wake up only when new matching mail or an upcoming calendar event appears. Unattended runs pause on world-changing actions and ask for approval in Telegram or the JOBS tab before continuing.
- Two channels, one brain. The browser UI and a Telegram bot share the same turn loop, so files, approvals and results flow both ways.
How we built it
- Agent runtime: Strands Agents SDK (TypeScript) on Amazon Bedrock AgentCore. Each session gets its own agent, streamed back as server-sent events so the room can animate every step.
- Compute: AgentCore Code Interpreter as the agent's sandbox, implemented by writing the SDK's abstract
Sandboxclass over the service's file and execution API. AgentCore Browser drives a real headless Chrome over the Chrome DevTools Protocol, with the SigV4 signature carried in the WebSocket URL so no extra libraries were needed. - Identity and data: Cognito for sign-in, DynamoDB for user records, agents, routines and jobs, S3 for files and chat snapshots, the AgentCore Identity vault for OAuth tokens, EventBridge Scheduler for routines.
- UI: a hand-written canvas renderer. The room is a 196x146 pixel buffer upscaled with no smoothing, drawn with an integer scanline filler so nothing antialiases into mush. Browser screenshots go on a second, full-resolution canvas mapped onto the monitor's parallelogram so the page stays readable. Three themes are pure data.
- Reliability rails: conversation summarisation, retry with backoff, per-invocation turn limits, a deadline around every tool, per-day token metering shown as an ENERGY bar, and a scripted eval suite.
Challenges we ran into
- AgentCore has no public URL. The runtime only exposes
/invocationsbehind SigV4, so the UI had to ship as a separate App Runner service that signs calls on the browser's behalf. Two architectures (arm64 for the agent, x86 for the UI), two images, two roles. - Streaming through every layer. If the runtime answers with JSON instead of an event stream, the whole room collapses to "thinking, then done". Getting the stream, including the inner events of delegated colleagues, attributed to the right desk took real work.
- Sandboxes bill while alive. An innocent file listing at the top of every turn was quietly booting a microVM for "what is 2+2". Sessions now start lazily and stop when the chat is evicted.
- Two filesystems confuse a model. Host-side file tools plus a sandbox left the agent unsure which disk it was on. We removed the host tools entirely.
- Isometric maths. Moving a character "toward the camera" means adding to both grid axes equally, and a monitor on the desk without a stand disappears behind the seated agent's head.
What we learned
Making the agents' work visible changed how we built the agents. Every time the room looked wrong, it was because the system was doing something wrong: booting VMs it did not need, delegating without context, reading files from the wrong disk. The pixel office turned out to be our best debugging tool.
We also learned that least privilege is easier to reason about when it is physical. A keycard on an agent's chest and a desk with no computer say more than a permissions table.
Built With
- amazon-bedrock
- amazon-bedrock-agentcore
- amazon-cognito
- amazon-dynamodb
- amazon-ecr
- amazon-eventbridge
- amazon-web-services
- anthropic
- aws-app-runner
- claude
- docker
- express.js
- gmail-api
- google-calendar-api
- html5
- javascript
- node.js
- strands-agents
- typescript
Log in or sign up for Devpost to join the conversation.