Inspiration

The TikTok TechJam Track 1 starter kit ships as a deliberately single-user Agent platform - any authenticated caller can view, edit, delete, or message any Agent. That's a real gap: in an actual multi-user platform, one user's coding agent, workspace, and protected data would be fully exposed to everyone else, and even a misdirected Agent could go looking for another user's files. We wanted to close that gap properly, not with a login screen that just hides things in the UI, but with real enforcement.

What it does

We built ownership isolation across three separate layers, so a failure at one layer doesn't expose the whole system:

  1. API layer — the backend rejects any request from a user who doesn't own the Agent they're trying to access, with a real 403 Forbidden.
  2. Infrastructure layer — each Agent's Docker container only has its owner's data mounted in. Even if something went wrong with the Agent's own reasoning, there's nothing belonging to another user physically present for it to find.
  3. Observability layer — every run and every denied access attempt gets logged automatically.

A simple User A / User B switcher in the sidebar lets you demo this live: create an Agent as User A, switch to User B, and watch that Agent disappear entirely; it's not hidden but genuinely filtered out server-side.

How we built it

  • Added an ownerId field to the Agent data model, required at creation.
  • Centralized the API-level ownership check in a single chokepoint - AgentService.getAgent(id, userId) - which every other method calls through first, instead of scattering checks across every route.
  • Methods that re-fetch the Agent inside an atomic store mutation re-check ownership inside that mutation too, closing a small gap between the initial check and the mutation actually running.
  • Built container-codex-runner.ts to dynamically construct a Docker --mount argument scoped to the calling Agent's owner, so the container's filesystem physically only contains that owner's data.
  • Built audit-logger.ts to write an append-only log of RUN_STARTED, RUN_COMPLETED, RUN_FAILED, and ACCESS_DENIED events, viewable in the UI.

Challenges we ran into

Getting the local dev environment running was its own challenge, from Docker setup to figuring out how to get a BytePlus ModelArk API key and endpoint. We also hit a real bug during development: the container mount logic looked correct in code, but wasn't actually taking effect during live runs. After ruling out Docker itself (tested the exact mount syntax directly and confirmed it worked), we traced it to a runtime configuration flag, RUNTIME_PROVIDER, defaulting to a code path that never applied the mount at all. That reinforced testing actual behavior, not just reading code and assuming it works.

Accomplishments we're proud of

  • Three layers of defense instead of one, so a bypass at any single layer doesn't expose user data.
  • 17 automated tests covering the ownership boundary at both the service and HTTP layers.
  • Verified everything live, via curl and direct API calls, not just the UI - this is what caught the container mount bug above.

What we learned

How much a single, well-placed enforcement chokepoint simplifies both implementation and testing compared to scattering authorization logic everywhere. Also that a feature that looks correct in code still needs to be verified running, since UI-only testing can mask real bugs underneath.

What's next

Real session-based authentication instead of the mock header we're using now, ownership checks on the one endpoint we know is still unguarded (GET /api/runs/:id), and a more structured, filterable view for the audit log.

Built With

+ 13 more
Share this project:

Updates

Submission history