V1

Inspiration

Inspiration

AI agents can be useful, but they may also follow unsafe instructions or call tools without sufficient verification. We want to build a simple safety layer that helps users understand and control an agent's risky actions.

What We Plan to Build

AgentGuard will inspect agent requests, identify possible prompt injection or risky tool calls, and ask for human approval before sensitive actions. The goal is to make AI agents safer and easier to trust.

How We Plan to Build It

We plan to build a lightweight web application with:

  • A prompt and tool-call inspection interface
  • Risk classification and explanation
  • Human approval checkpoints
  • A simple audit log
  • A demo agent that performs safe, simulated actions

Challenges

The main challenges will be reducing false alarms, explaining risks clearly, and designing approval steps that improve safety without making the agent difficult to use.

What We Hope to Learn

We hope to learn how to evaluate agent safety in practical workflows and how open AI infrastructure can support reliable, transparent agent systems.

Important Note

This project is a new hackathon project. The existing Agent Canary repository is used only as background experience and inspiration, not submitted as the completed hackathon project.

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for AgentGuard

V2 Before an AI agent acts, AgentGuard checks the risk. AgentGuard is a lightweight safety layer for AI agents. It inspects proposed prompts and tool calls before execution, detects prompt injection, sensitive-data exposure, and high-impact actions, then explains the risk in plain language. Users can simulate agent actions, review matched policies, choose a stricter sensitivity mode, and approve or reject risky operations. All decisions stay in the browser, making the demo fast, private, and easy to understand.

Prompt injection detection / 提示词注入检测 Sensitive data protection / 敏感数据保护 High-impact action review / 高影响操作审批 Risk sensitivity modes / 风险敏感度模式 Agent tool-call simulation / Agent 工具调用模拟 Policy match explanation / 策略命中解释 Event history and export / 事件历史与报告导出 Bilingual interface / 中英文界面

AgentGuard is a web-based safety layer for AI agents. It simulates browser-shield behavior by inspecting prompts and proposed tool calls, detecting prompt injection, sensitive-data exposure, and high-impact actions before execution.

Built With

Share this project:

Updates

Submission history