Inspiration

AI agents are increasingly able to modify files, update configurations, query private data, and call external services. Traditional role-based access control answers one important question: “Is this agent allowed to perform this action?” But permission alone does not make an action safe. An agent may be authorised to update a deployment configuration while behaving unusually, affecting sensitive systems several dependencies away, or acting outside its established operational pattern. By the time a dashboard reports the problem, the side effect may have already happened. We built QuantQueens Agent Safety Middleware to intervene before that point. It acts as a runtime “bouncer” between an AI agent and a protected resource.

What it does

Before a managed action can execute, QuantQueens verifies: Identity: Which human, agent, and run initiated the request? Authorisation: Does the agent have the exact capability for this resource? Ownership: Is the resource owned by the correct principal? Graph impact: What downstream services or sensitive datasets could be affected? Behavioural context: Does the action differ from trusted, successful previous runs? Circuit-breaker state: Should the action proceed, pause for review, or be blocked? Only an allowed action receives a single-use execution claim. Denied, risky, or circuit-breaker-tripped actions stop before the protected adapter changes anything.

Our primary demo shows an Alice owned agent successfully reading Alice’s managed record while being denied access to Bob’s record, even when the caller attempts to forge Bob’s identity.

The second scenario demonstrates why the project goes beyond RBAC. An agent retains valid permission to modify a production configuration, but the middleware discovers that the action is historically novel and could affect five downstream resources, including customer data. Authorization returns ALLOW, contextual risk returns BLOCK, and the production configuration remains unchanged.

How we built it

QuantQueens extends the Volc Agent Launchpad starter with two enforcement layers: A coarse pre-run gate that can pause or deny an agent run before Codex starts. A stronger Resource Gateway immediately before managed resource side effects.

The backend is built with Node.js, TypeScript, Fastify, Zod, and SQLite.

It maintains: Server-attested human, agent, run, and delegation identities Exact capabilities and resource ownership A backend-queryable knowledge graph Deterministic downstream-impact analysis Trusted-history behavioural baselines Persistent NORMAL, WARN, and TRIPPED circuit-breaker states Atomic, single-use execution claims Idempotent effect receipts Structured run events with strictly increasing sequence numbers

Delegation cannot create new authority. Its effective capability is the safe intersection. The React and Vite interface turns this evidence into plain-language decisions, impact maps, network graphs, and an ordered audit timeline that survives reloads and restarts. General agent tasks can run through Codex CLI using the Volcengine Ark Responses API. The deterministic security demonstration does not require model credentials.

Challenges we faced

The hardest challenge was proving that an action was genuinely prevented, not merely displaying a warning after it happened. We designed a narrow gateway in which the managed SQLite adapter cannot execute without a valid, atomically claimed decision. Tests verify that blocked actions leave both the adapter invocation count and durable resource state unchanged. Another challenge was keeping authorisation and behavioural risk separate. Graph proximity, historical behaviour, and ownership can increase or decrease risk, but they must never grant a missing permission.

We also had to make historical behaviour resistant to poisoning. Denied, blocked, failed, and unconfirmed prompt only actions do not become part of the trusted normal baseline, no matter how often they are attempted.

Finally, distributed audit events are difficult to order reliably. Timestamps alone are insufficient, so every event receives an atomically allocated, run local sequence number. This lets us reconstruct each request, authorisation decision, risk result, circuit breaker transition, attempted effect, and final outcome deterministically.

What we learned

We learned that trustworthy agent systems require more than prompts, static policies, and after-the-fact logs. Effective enforcement needs a trusted runtime boundary connecting identity, exact authority, dependency context, historical evidence, and the real side effect.

We also learned that a knowledge graph becomes significantly more useful when it participates directly in backend policy decisions. In QuantQueens, the graph is not decorative visualisation, it helps determine blast radius before execution.

Most importantly, explainability must be built into the decision itself. Every result records who attempted what, which resource was involved, why the risk changed, what would have been affected, and whether the side effect actually occurred.

Current scope and future work

QuantQueens is an intentionally focused proof of concept. Its strongest guarantees currently apply to resources routed through the managed Resource Gateway; it does not yet transparently mediate every arbitrary shell, filesystem, or network operation performed by Codex. Next, we would add additional protected adapters, integrate production identity and multi tenant sessions, and introduce a transactional outbox for reliable external side effects.

Built With

Share this project:

Updates

Submission history