V1
Inspiration
Inspiration
AI agents can be useful, but they may also follow unsafe instructions or call tools without sufficient verification. We want to build a simple safety layer that helps users understand and control an agent's risky actions.
What We Plan to Build
AgentGuard will inspect agent requests, identify possible prompt injection or risky tool calls, and ask for human approval before sensitive actions. The goal is to make AI agents safer and easier to trust.
How We Plan to Build It
We plan to build a lightweight web application with:
- A prompt and tool-call inspection interface
- Risk classification and explanation
- Human approval checkpoints
- A simple audit log
- A demo agent that performs safe, simulated actions
Challenges
The main challenges will be reducing false alarms, explaining risks clearly, and designing approval steps that improve safety without making the agent difficult to use.
What We Hope to Learn
We hope to learn how to evaluate agent safety in practical workflows and how open AI infrastructure can support reliable, transparent agent systems.
Important Note
This project is a new hackathon project. The existing Agent Canary repository is used only as background experience and inspiration, not submitted as the completed hackathon project.
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for AgentGuard
V2 Before an AI agent acts, AgentGuard checks the risk. AgentGuard is a lightweight safety layer for AI agents. It inspects proposed prompts and tool calls before execution, detects prompt injection, sensitive-data exposure, and high-impact actions, then explains the risk in plain language. Users can simulate agent actions, review matched policies, choose a stricter sensitivity mode, and approve or reject risky operations. All decisions stay in the browser, making the demo fast, private, and easy to understand.
Prompt injection detection / 提示词注入检测 Sensitive data protection / 敏感数据保护 High-impact action review / 高影响操作审批 Risk sensitivity modes / 风险敏感度模式 Agent tool-call simulation / Agent 工具调用模拟 Policy match explanation / 策略命中解释 Event history and export / 事件历史与报告导出 Bilingual interface / 中英文界面
AgentGuard is a web-based safety layer for AI agents. It simulates browser-shield behavior by inspecting prompts and proposed tool calls, detecting prompt injection, sensitive-data exposure, and high-impact actions before execution.
Log in or sign up for Devpost to join the conversation.