Inspiration

AI agents are becoming capable of working on real codebases, running commands and editing files with very little supervision. That is useful, but it also raises a basic question: what happens when an Agent makes the wrong change, consumes too many resources or refuses to stop?

We wanted to explore the infrastructure around the model rather than build another chatbot interface. Our goal was to make Agent execution safer and more controllable without relying on the model to correctly report its own actions.

That led us to build TechTokers, a governance layer for the Agent Launchpad starter kit.

What it does

Our system governs an Agent at three different stages.

Before execution: Resource Admission

Each Agent can have optional limits for:

  • Maximum prompt size
  • Maximum number of admitted Runs
  • Maximum cumulative model-token usage

These checks happen in the trusted backend before the Runtime is invoked. If a request exceeds its policy, it is rejected without calling the model or consuming another Run slot.

The admission decision and actual token usage are persisted as redacted governance evidence.

During execution: Runtime Containment

Admitted Runs execute through Codex CLI inside disposable Docker containers.

Per-Agent runtime policies can limit:

  • Execution duration
  • Output size
  • CPU usage
  • Memory usage
  • Process count

The user can stop a Run normally or kill it when containment is required. Repeated runtime-limit violations can automatically quarantine an Agent.

After execution: Transactional Workspace Protection

The Agent never edits its persistent workspace directly.

For every Run, the backend creates an isolated staging copy and mounts that copy into the container at /workspace. Files such as .env and .env.* are excluded before the copy is created, so they are not merely hidden in the interface—they are absent from the Agent's working environment.

After the Run, the backend independently examines the staging filesystem and generates a SHA-256 manifest of created, modified and deleted files. The model's textual description of its work is not trusted as evidence.

The resulting changes follow one of two modes:

  • Review mode: every change set requires human approval.
  • Auto mode: ordinary source and documentation changes may be applied automatically, while risky changes are escalated and protected changes are denied.

Approved changes are applied using a rollback-capable transaction engine. Conflicting, denied, failed or expired proposals do not modify the persistent workspace.

How we built it

We extended the provided React and Fastify application while keeping its existing AgentService, AgentRunner, JSON persistence, Codex sessions and ModelArk integration.

The trusted execution path is roughly:

  1. The React Playground sends a task to Fastify.
  2. AgentService evaluates the Agent's resource policy.
  3. The backend creates a protected per-Run staging workspace.
  4. AgentRunner launches Codex CLI inside a constrained Docker container.
  5. Codex communicates with the ModelArk Responses API.
  6. The backend measures actual usage and inspects the resulting filesystem.
  7. A verified change set is classified as automatic, review-required or denied.
  8. Approved changes are transactionally applied to the persistent workspace.
  9. State, decisions and redacted evidence are persisted for the UI.

The important design decision was to keep enforcement outside the model. Prompting the model differently cannot disable admission checks, container limits or change-set classification because those decisions are made by backend code using persisted state and observed filesystem data.

Challenges we faced

The hardest part was deciding where a trustworthy control boundary could actually exist.

Our first idea was an MCP-based protected-resource gateway. Codex CLI could register our MCP server, but our configured ModelArk model path did not produce the structured tool calls required for a real protocol-native round trip. We could have parsed textual JSON from the model, but that would only imitate enforcement while remaining bypassable. We abandoned that approach instead.

This pushed us toward a provider-independent design: allow Codex to work normally inside staging, then treat every filesystem change as an untrusted proposal.

Transactional application was another challenge. Applying several files is easy when everything succeeds; the difficult case is when the third operation fails after the first two have already changed the workspace. We introduced validation, base-hash conflict detection, transaction journals, backups, atomic replacement, deferred deletion and rollback verification to prevent partial application.

We also had to be precise about what our system does not guarantee. Transactional workspace protection controls what becomes persistent, but it does not classify every Bash command or reverse external network side effects. Runtime containment limits the blast radius, but it is not a command-level authorization engine.

What we learned

The main lesson was that reliable Agent infrastructure cannot be built entirely through prompts.

The model can explain what it believes it changed, but the filesystem is the source of truth. A UI warning is not protection if the Runtime can still access the underlying file. An approval button is not meaningful unless the change remains isolated until approval. Similarly, a kill button is useful for containment, but it does not replace resource limits or transactional persistence.

We also learned that limitations are worth documenting honestly. This is a single-process proof of concept using JSON-backed persistence. Its atomic admission, active-Run coordination and expiry reconciliation are not distributed guarantees. It does not yet provide user roles, a multi-node transaction coordinator, an external immutable audit trail or command-level network policy.

Even with those limits, the project demonstrates a practical principle: Agents may propose work, but trusted infrastructure should decide what is admitted, what is allowed to continue and what is permitted to persist.

Built With

Share this project:

Updates

Submission history