TrustLayer AI—Securing Every AI Decision Before It Becomes an Action
Track: AI + Cybersecurity
Inspiration
Artificial intelligence is moving beyond answering questions. AI agents can now read emails, access documents, communicate with other services, and make decisions on behalf of users.
But what happens when an AI agent encounters a malicious instruction disguised as a legitimate email, document, or business request? An attacker might impersonate a manager, request a fraudulent payment, or hide instructions inside an external document to trick an AI agent into sharing confidential information.
We built TrustLayer AI to address this problem. Our goal is to make AI agents safer by introducing a security checkpoint between the actions they propose and the actions they are allowed to perform.
What It Does: TrustLayer AI is an AI-agent security platform that evaluates proposed actions before execution. It analyzes the source of an instruction, the action an agent wants to perform, the resources involved, the sensitivity of the information, and the permissions granted to the agent.
TrustLayer then produces one of three security decisions: ALLOW: The action meets the configured security requirements. REVIEW REQUIRED: The action requires additional human approval. BLOCK: The action violates a security policy or presents an unacceptable risk.
Each analysis includes security reasoning and risk indicators to help users understand the decision. The platform also provides a security dashboard, interactive scenarios, policy visibility, and audit logging.
How We Built It: TrustLayer combines a React-based frontend with a Python FastAPI backend.
Its security architecture includes:
- Risk analysis: Evaluates suspicious instructions, untrusted sources, sensitive data access, and potentially dangerous actions.
- Policy enforcement: Checks whether an AI agent has permission to perform a requested operation.
- Contextual analysis: Supports AI-assisted evaluation of suspicious instructions, with deterministic security rules remaining authoritative.
- Decision engine: Combines security findings into an explainable ALLOW, REVIEW REQUIRED, or BLOCK decision.
- Audit logging: Records security decisions for review and investigation.
- Interactive dashboard: Makes the security workflow accessible through a web interface.
The prototype uses simulated AI-agent scenarios to demonstrate how these protections can work.
Challenges We Faced: One of our biggest challenges was designing a security system that could distinguish between legitimate user requests and potentially malicious instructions embedded in external content. We also had to balance security with usability. Blocking every action would make AI agents ineffective, while allowing everything would expose users to serious risks. Another challenge was connecting the frontend, backend, policy engine, and audit system into a consistent workflow while preparing the application for deployment. These challenges shaped our decision to prioritize explainable security decisions and a clear separation between trusted user instructions and untrusted external content.
Accomplishments We're Proud Of: We built a security-focused prototype that demonstrates how proposed AI-agent actions can be evaluated before execution. TrustLayer brings together risk scoring, permission checks, decision explanations, and a security dashboard in one application. Rather than treating AI output as automatically trustworthy, the system evaluates what an agent is attempting to do and why that action may be risky.
What We Learned: Building TrustLayer reinforced an important cybersecurity principle: an AI agent should not receive unlimited authority simply because it is acting on behalf of a legitimate user. We learned more about indirect prompt injection, least-privilege permissions, risk-based decisions, API integration, and the importance of maintaining audit records. We also learned that AI-assisted security analysis is most useful when combined with clear, enforceable security policies.
What's Next: Our long-term goal is to extend TrustLayer beyond simulated scenarios into real AI-agent workflows. Future improvements include integrations with agent frameworks, stronger identity and permission controls, production-grade database persistence, security analytics, and human approval workflows. We envision TrustLayer becoming a security layer that helps organizations deploy autonomous AI systems with greater visibility, accountability, and control.
Built With: Python, FastAPI, React, TypeScript, Vite, SQLite, REST APIs, and AI-assisted contextual security analysis.
Links: Live Demo: https://trustlayersecurity.vercel.app Source Code: https://github.com/Benedict-Adjei/trust-layer.git
Demo Video:
Current Limitations: TrustLayer is a prototype that evaluates simulated agent actions. It does not independently intercept arbitrary third-party AI agents or execute external actions. Production enforcement would require integration into an agent's execution workflow. Deployment-dependent capabilities, including contextual AI analysis and persistent audit storage, should be evaluated against the live environment before production use.
Built With
- fastapi
- featherless-ai
- git
- github
- postgresql
- pydantic
- python
- react
- sqlalchemy
- supabase
- tailwindcss
- typescript
- vercel
- vite
Log in or sign up for Devpost to join the conversation.