Inspiration
Organizations and governments are racing to put AI agents to work, including writing code, touching files, calling tools, accessing data, and identity is the part nobody's solving. The starter kit we built on is a real example of the default state most agent platforms ship in: any user can view or modify any agent, any agent can touch any workspace or data, there's no record connecting a human's request to an agent's action, and no way to stop a misbehaving agent mid-task. That's fine for a solo demo. It falls apart the moment more than one person shares the platform, which is exactly what happens the moment an organization deploys something like this for real. We wanted to prove that "agents get capabilities, not credentials" could be a real, testable architecture, not just a slogan.
What it does
Agent Launchpad is a middleware layer for the Identity & Authorisation track that sits between a human, an AI agent, and the tools/data that the agent can touch:
- Every agent runs and mints its own scoped, revocable identity: An agent JWT plus a RunToken store row (the authoritative source, never the cached JWT claims), narrowed to exactly the five tool scopes it holds and further auto-narrowed to what the current task actually needs, before the model ever runs, which is inspired by Progent (arXiv 2504.11703).
- Every tool call is checked at a gateway: Who, which tool, whose data, where it's allowed to go, logged as a RunEvent on every branch, allow and deny alike.
- A human confirms every grant via an Access Request Card (Allow for this run / Always allow / Deny): Deny-by-default, and nothing widens access without a human seeing it.
- Information-flow control tracks not just where data came from but whether it can be trusted: Content read from a borrowed workspace carries a trust label separate from its sensitivity label, so a prompt-injected instruction can't turn a read grant into an exfiltration channel, even if it leaks nothing an attacker could point to directly (inspired by FIDES, arXiv 2505.23643).
- Two isolation layers, enforced independently: Intra-user (your own agents don't trust each other by default) and multi-tenant (Postgres row-level security, owner-only, enforced even if the gateway is bypassed).
- Grants are revocable mid-task: Pulling one grant 403s only that call going forward; the rest of the agent's work keeps running.
How we built it
TypeScript across a Fastify control plane and a React UI, with Codex CLI running in a disposable Docker container per agent turn. Identity tokens are signed/verified with jose (HS256). The gateway is exposed over MCP streamable HTTP (@modelcontextprotocol/sdk) at POST /mcp, with a documented REST fallback. Postgres holds the one protected resource that needs real row-level security (crm_records); everything else (agents, runs, tokens, grants, events) stays in the existing JSON store, so the baseline app never had to be rearchitected. We verified early, via a deliberate spike, that Codex's exec command actually sends the Authorisation header from -c mcp_servers.*.http_headers on initialisation, which is the one assumption the whole per-run-token design depends on.
Challenges we ran into
The permission narrowing feature almost silently broke our own demo. Once we landed automatic scope narrowing (only granting a run the tools its prompt actually needs), one of our injection-attack demo scenes stopped triggering at all, the narrowed token no longer held the scope the attack needed, so there was no deny, no card, and no event to show. It wasn't a bug; it was the feature working exactly as designed. We had to consciously decide how to rehearse around it rather than quietly patch it away.
Also, ordering between our two IFC checks was load-bearing, not incidental. Confidentiality (can this data leave at all) has to run before integrity (can this content be trusted) inside the same gate, or the demo's centrepiece deny reads as the wrong reason. Docker networking wasn't portable by default because host.docker.internal didn't resolve the same way across Docker on Mac/Windows vs. Podman on Linux, which needed explicit --add-host handling.
Fitting six demo scenes into a three-minute window with only two genuinely live runs meant deciding upfront which moments had to be real and which could fall back to a scripted/demo/replay that still fires real gateway calls with a real token, rather than a recording.
Accomplishments that we're proud of
We treated information-flow control as a committed feature from day one, not a stretch goal, where we assume most teams would cut it first when time runs short. Automatic, task-scoped permission narrowing means an agent's tool menu shrinks before the model even sees the prompt, so a planted instruction literally cannot reach a tool the task never asked for, not just "denied," but never offered.
Also, row-level security holds even if our own gateway has a bug and lets a call through. Defence in depth we could actually prove with a test, not just claim. Revoking access mid-task narrows what an agent can still do without killing the whole run, accountability doesn't have to mean throwing away in-progress work.
What we learned
The direction of a permission change matters more than the mechanism. Our clearest design rule became: removing permissions can happen automatically and needs no human, but adding one always needs a card. That one asymmetry made most of the rest of the system's behaviour predictable to reason about.
Also, never ask a model whether content is trustworthy. A model reading attacker-controlled text can simply be told to lie about it. Trust has to be decided by how content arrived (your own workspace vs. a borrowed one), never by asking about the thing that's being read, potentially poisoned input. Another thing is that a feature can be correct and still break your demo. The scope-narrowing incident taught us to test security features against our own attack scenarios, not just against unit tests. "Does this still show what we want the audience to see?" is a real requirement.
What's next for #openToWork
Our next goal is to develop a platform-admin view with cross-agent visibility and the ability to intervene, even on agents the admin doesn't own, alongside cross-tenant grants with explicit double opt-in so sharing can happen between organisations, not just within one. We also want to move from run-level taints to per-value provenance, tracking what a specific fact touched, inspired by CaMeL, so a single CRM lookup doesn't blanket-block the rest of an unrelated task, and to bundle repeat approvals into a single decision instead of a card per denial, cutting interruption fatigue without weakening the guarantee. Finally, we plan to swap in a real identity provider over OIDC/SAML; the seeded, no-password login was a deliberate scope cut for the hackathon, and the whole system was built so that swap touches one file, auth.ts, and nothing else.
Hope you learnt something from our project, till next time! Ciao!
Built With
- docker
- fastify
- node.js
- postgresql
- react
- typescript
- vite
Log in or sign up for Devpost to join the conversation.