Inspiration
AI agents are moving beyond answering questions. They can now run shell commands, modify files, access credentials, install software, and interact with real infrastructure.
Most AI safety tools still focus on the prompt: they inspect what a user says but have limited visibility into what the agent actually does afterward. A prompt can look harmless while containing instructions that cause the agent to leak data or misuse a tool.
We built Sentinel to place security at the action boundary. Instead of relying only on the model to behave safely, Sentinel independently monitors and governs the commands it executes.
What it does
Sentinel is a real-time governance layer for autonomous AI agents. Every command reported by the agent runtime is evaluated against policy and assigned one of five outcomes:
- Allow safe actions to continue.
- Hold uncertain actions for verified human approval.
- Block serious violations, such as attempts to access protected secrets.
- Terminate runs that exceed their command budget or behave unexpectedly.
- Recover safely so the agent remains usable after containment.
Risky but potentially legitimate actions can be reviewed by a human. Each reviewer uses a personal access token, so Sentinel automatically records who approved or denied the request. The approver’s identity cannot simply be typed or impersonated.
Sentinel also stores a redacted audit trail and displays security results through a live evaluation dashboard.
How we built it
We extended the CodeJam starter kit using React, Fastify, TypeScript, Codex CLI, and disposable Docker, Colima, or Podman containers.
Codex produces a streamed JSON event whenever it starts a command. Sentinel evaluates that event through three layers:
- A shell parser identifies commands, tools, destinations, and protected resources.
- A capability engine determines what the command is attempting to do, such as reading a secret or accessing the network.
- A policy engine decides whether to allow, hold, block, or terminate the action.
Each agent run executes inside a disposable container. With network isolation enabled, the agent has no direct outbound route and can only reach the model endpoint through a dedicated broker.
The React interface displays held actions, verified approval controls, policy decisions, recovery status, and benchmark results.
We also created a one-command local startup flow:
npm run poc
Challenges we ran into
The biggest challenge was that shell commands can express the same action in many different ways. A simple network request can be written using curl, Python, Node.js, shell redirection, command substitution, encoded text, or a custom script.
Our first hand-written test corpus did not cover every form. We created a generated attack bank to combine different tools, destinations, encodings, and shell wrappers. Each time we expanded its axes, it revealed bypasses that the original tests had missed.
We also learned that command-level enforcement is reactive. Sentinel receives the command while it is executing, so a very fast command may produce a partial effect before the container is destroyed. This led us to add network-level isolation rather than relying entirely on command parsing.
Other challenges included:
- Keeping model credentials available to Codex but hidden from spawned shell commands
- Resuming held runs without creating permanent allowlist changes
- Redacting sensitive information from audit evidence
- Preventing runaway agents without limiting normal work
- Ensuring our security claims accurately matched what we measured
Accomplishments that we're proud of
We built a complete governance loop where every outcome—allow, hold, block, terminate, and recover—is enforced by the backend rather than simulated in the interface.
On our authored evaluation corpus, Sentinel achieved:
- 0 of 114 attacks allowed
- 0 of 40 secret-channel attacks allowed
- 1 of 84 legitimate tasks blocked
- All 6,860 generated attack variants detected
These are policy results on our test sets, not a claim that Sentinel can prevent every possible real-world attack. We are equally proud that the project reports that limitation clearly.
We also:
- Demonstrated container-level network isolation using a live mock collector
- Implemented run-scoped approvals tied to verified reviewer identities
- Preserved agent recovery after containment
- Built a live security evaluation dashboard
- Created a one-command startup experience
What we learned
We learned that no single security layer is sufficient.
The model may refuse a malicious request, but model behavior is not guaranteed. A command policy can catch known attack patterns, but parsers have blind spots. Network isolation can prevent outbound traffic, but it cannot explain why an action was dangerous. Human approval adds context, but only when reviewer identity and decisions are properly recorded.
Sentinel therefore combines:
- Model behavior
- Command policy
- Network containment
- Human approval
- Resource limits
- Redacted audit evidence
We also learned that security benchmarks are only as useful as their scope. A perfect score may simply mean the test set cannot express the missing attack. Independent review and honest documentation are just as important as the headline number.
What's next for Sentinel
Sentinel currently focuses primarily on shell commands. Our next step is to extend the same capability-based governance model to other agent actions, including API requests, MCP tool calls, database operations, and cloud infrastructure changes.
We also plan to add:
- Role-based approval permissions
- Token rotation and revocation
- Multi-user and multi-tenant isolation
- Durable append-only audit storage
- More configurable organizational policies
- Independent red-team datasets and external evaluation
- Production-grade deployment and monitoring
Our long-term goal is to let organizations adopt autonomous agents without surrendering control over their data, systems, or infrastructure.
Built With
- byteplus
- docker
- esbuild
- fastify
- github-actions
- node.js
- openai-codex
- podman
- react
- terraform
- typescript
- vite
- vitest
- volcengine
- zod

Log in or sign up for Devpost to join the conversation.