Inspiration
AI agents can make decisions using outdated, incomplete, or incorrect information. A normal AI safety check may ask another AI to review the decision, but that second AI can make the same mistake.
We wanted to build a system that checks an AI agent's intended actions against the company's actual data and policies before anything happens.
What it does
ControlPlane is a runtime safety layer for AI agents.
Before an AI agent executes a tool call, ControlPlane:
- extracts the claims behind the action,
- checks them against fresh company data,
- evaluates the relevant business rules,
- decides whether to ALLOW, BLOCK, MODIFY, or ESCALATE the action,
- creates a signed receipt explaining the decision.
For example, an AI agent can try to issue a ₹42,999 refund based on outdated information. ControlPlane checks the actual order and the current refund policy, and blocks the refund before it executes.
How we built it
We built ControlPlane in Python with an AI agent, structured claim extraction, fresh evidence from company databases, rule-based policy evaluation, and signed decision receipts.
The system uses an interception layer so the same verification engine can protect different AI agents without changing the agent's core code.
We also built a FastAPI dashboard so judges can run realistic scenarios and see the evidence, decision, intervention, and execution result.
Challenges we ran into
The biggest challenge was avoiding another AI-based opinion and instead grounding decisions in trusted company systems.
We also had to make the engine independent of any specific use case, keep evidence and decisions auditable, and ensure the demo was reproducible without requiring external API keys.
Accomplishments that we're proud of
We built a single verification engine that supports multiple use cases, including refunds, document access, and discount approval.
We also built an eight-scenario judge dashboard, signed decision receipts, automated tests, and a reproducible offline demo.
On our internal 140-case gold set, record-grounded verification achieved 100.0% direction accuracy compared with 87.9% when using the agent's own retrieved evidence. This result is specific to our evaluated dataset and failure mode.
What we learned
We learned that reliable AI systems need more than better prompts or another LLM checking the first LLM.
The most important question is often: "What does the real system say right now?"
We also learned that AI governance should be measurable, reproducible, and transparent about its limitations.
What's next for ControlPlane
Next, we want to validate ControlPlane on larger external benchmarks, support more enterprise systems, improve human review for uncertain decisions, and make it easier to add the verification layer to existing AI agents.
Built With
- agents
- ai
- engine
- fastapi
- generative
- governance
- hmac-sha256
- learning
- llm
- machine
- policy
- python
- qwen
- rag
- safety
Log in or sign up for Devpost to join the conversation.