Inspiration

AI agents are getting powerful rapidly, they can send emails, modify databases, score job candidates, and take actions that affect real people. But there's a major caveat. We kept asking: what happens when an agent does something it shouldn't? What if it sends a rejection email to every applicant, or applies a biased scoring rubric without anyone noticing?

In 2015, Amazon penalized points against those who used the word "women" in their resumes. Recent research finds that a vast majority of Anthropic developers find their enterprise agents acting out of scope. These are big names that create products that affect millions of people globally.

This is where SafeAgent steps in. We wanted to build the safety layer that sits between an AI agent's intention and action, something that catches misalignment before it causes harm, not after.


What it does

SafeAgent is a real-time safety gate for AI agent systems. Before any agent executes a tool call, SafeAgent intercepts it and runs it through a three-tier pipeline:

  • T1 Guardrails: instant deterministic blocklist (~0ms). Blocked tools like bulk_delete or send_to_all never make it through.
  • T2 Semantic Cache: lookup on repeated patterns (~5ms). If we've seen this action before, we skip Claude entirely.
  • T3 Claude Sonnet 4.6 Scoring: structured AI scoring of misalignment (0–100) and oversight risk (0–100), with an auto-generated safer alternative (~800ms).

Every action gets one of three verdicts:

  • ALLOW → agent executes normally
  • WARN → human reviews Claude's safer alternative before proceeding
  • BLOCK → hard stop, no recovery

On top of the safety gate, Fetch.ai acts as an intelligent agent recommendation system, analyzing the builder's intent and recommending the right agents for the job before the graph is even built.

Every gate event is logged and exported as safe-agent-blueprint.json for full auditability.

How we built it

We split ownership across three layers that had to talk to each other perfectly:

Joseph built the LangGraph orchestration layer using Anthropic Claude, a meta-scaffolder that takes a plain-English builder description and generates a full multi-agent graph with tool assignments, model tiers (Haiku for classification, Sonnet for reasoning), and A/B topology selection with cost prediction. Fetch.ai plugs in here as a recommendation engine, suggesting the best agents for a given task based on the builder's intent.

Evan built the FastAPI Safety Gate — the T1/T2/T3 intercept pipeline and the safe-agent-blueprint.json audit export.

Utkarsh built the Vercel frontend with React Flow for live agent graph visualization, a flag modal for human-in-the-loop approve/modify/override decisions, an Arize Phoenix observability panel for A/B topology comparison, safety score drift tracking, and a Prediction vs Reality proof panel.

Challenges we ran into

Contract alignment was the hardest problem. Three people building independently means three sets of assumptions. The GateResponse expected decision as uppercase "ALLOW"/"WARN"/"BLOCK" and tier_triggered as an integer 1/2/3 , differences that would have been silent runtime failures.

Integrating Fetch.ai as a recommendation layer required mapping its agent discovery output to LangGraph's scaffolding input, two systems with very different mental models of what an "agent" is.

SSE streaming to Vercel required explicit X-Accel-Buffering: no headers and keepalive pings to stop proxies from dropping the live event connection after 30 seconds of low traffic.


Accomplishments that we're proud of

  • A three-tier gate pipeline that goes from 0ms (T1) to ~800ms (T3) and short-circuits intelligently so that Claude only gets called when it needs to.
  • Fetch.ai agent recommendations that make the system smarter before execution even begins. Right agents, right task, every time.
  • Full auditability: very gate event, human decision, and override is captured and exportable as a blueprint JSON.
  • Real-time human-in-the-loop: the frontend lights up the moment a flag fires, shows scores, explanation, and Claude's safer alternative, and lets a human approve, modify, or override in seconds.
  • A working multi-service system: across three developers, three tech stacks; all within 24-hour hackathon.

What we learned

  • Agree on contracts before writing code - The interface between services is more important than the implementation inside them.
  • Claude's structured JSON output is almost production-ready when the system prompt is tight. We used it as a real-time scoring engine with consistent, parseable results.
  • Fetch.ai opens up a new paradigm - instead of hardcoding which agents to use, you can let a decentralized network recommend the best ones dynamically based on intent.
  • Human-in-the-loop is a product feature, not just a safety feature - Users trust a system more when they can see why something was flagged and choose what happens next.

What's next for SafeAgent

  • Deeper Fetch.ai integration - moving beyond recommendations to letting Fetch.ai agents autonomously register, discover, and compose themselves into pipelines based on a builder's intent.
  • Multi-tenant support - per-organization blocked tool lists, custom risk thresholds, and role-based override permissions.
  • Drift detection - alert when an agent's average safety score trends upward over a session, catching gradual misalignment before it becomes a problem.
  • SDK wrapper - a one-line Python decorator that wraps any LangChain or LangGraph tool call with SafeAgent protection, no manual API calls needed.
  • Fine-tuned scoring model - replace general Claude scoring with a model fine-tuned specifically on agent safety decisions, dropping T3 latency from ~800ms to ~100ms.

Built With

  • anthropic-claude-api
  • arize-phoenix
  • asi
  • claude-haiku-4.5
  • claude-sonnet-4.6
  • docker
  • events
  • fastapi
  • fetch.ai
  • langgraph
  • next.js
  • python
  • react
  • react-flow
  • redis
  • server-sent
  • upstash
Share this project:

Updates

Submission history