Inspiration
Autonomous AI agents are moving from conversational systems to systems that can reason, call tools, delegate work, and take actions on behalf of users and organizations.
That creates a security problem that is fundamentally different from traditional application security.
An agent may begin with a legitimate task, but its reasoning context can contain untrusted information: a supplier document, an email, a webpage, retrieved RAG content, or another agent's output. If that content contains a prompt injection or malicious instruction, the problem is not simply that one model produced a bad response.
The more dangerous failure is propagation.
In a multi-agent workflow, one compromised agent can produce output that becomes trusted context for another agent. That second agent can pass contaminated reasoning to a third agent, eventually turning an untrusted instruction into an apparently legitimate tool action.
We were inspired by a simple question:
What if we treated an autonomous-agent failure like a security incident rather than a bad model response?
That led to AgentRescue.
AgentRescue is designed as a control plane for autonomous-agent reliability and security. Instead of assuming that an agent will always reason correctly, it assumes that failures can happen and builds mechanisms to detect, contain, investigate, verify, authorize, and remember them.
The core principle is:
An agent can recommend an action. Recommendation is not authorization.
That principle shaped the entire architecture.
What it does
AgentRescue demonstrates how a multi-agent system can respond when untrusted information contaminates an autonomous workflow.
The demonstration starts with a controlled supplier-document scenario.
A four-stage workflow contains:
Research Agent → Risk Agent → Procurement Agent → Purchase Action
A poisoned input introduces a prompt injection into the workflow.
Rather than allowing the failure to silently propagate, AgentRescue turns it into an explicit incident.
1. Detects and contains the failure
The system monitors the agent workflow and evaluates proposed tool actions through a zero-trust gateway.
The gateway does not blindly trust an agent because the agent is part of the system. It evaluates the action against identity, policy, scope, and the current security state.
When the simulated procurement action crosses the protected boundary, the gateway can BLOCK it before execution.
The important distinction is that containment happens at the action boundary — not after an unsafe action has already occurred.
2. Calculates blast radius
Once an incident is detected, AgentRescue reconstructs the propagation path through the agent graph.
For the demonstration incident, the system identifies how contamination moved between agents and calculates the resulting exposure.
This transforms:
"The model behaved incorrectly."
into:
"This input entered here, propagated through these agents, reached this proposed action, and was stopped at this boundary."
That makes the failure operationally understandable.
3. Launches a Google ADK investigation
AgentRescue uses Google Agent Development Kit (ADK) to orchestrate an investigation team.
The investigation is intentionally separated from execution. The investigation agents are given read-oriented capabilities to analyze evidence such as:
- prompt-injection indicators
- propagation paths
- blast radius
- compliance implications
- identity metadata
- institutional memory
They investigate the incident without receiving authority to execute the sensitive enterprise action.
This separation is central to the design:
Investigation ≠ authorization.
4. Verifies remediation instead of trusting it
A proposed remediation is not automatically considered safe.
AgentRescue provides a Remediation / Verification Lab that evaluates the proposed fix against adversarial mutations, regression scenarios, and compliance checks.
The demonstration executes a 100-test verification matrix:
| Category | Tests |
|---|---|
| Adversarial cases | 50 |
| Regression checks | 25 |
| Compliance checks | 25 |
| Total | 100 |
A remediation therefore has to produce measurable evidence before it can move toward recovery.
The goal is simple:
A patch is a proposal. Verification is evidence.
5. Requires explicit human authorization
Even after successful verification, AgentRescue does not automatically restore a sensitive financial workflow.
The system enters an approval state. A human must explicitly authorize deployment.
This creates a deliberate security boundary between:
AI recommendation → verified remediation → human authorization → recovery
rather than:
AI recommendation → automatic execution
6. Creates persistent institutional memory
AgentRescue includes Failure DNA and an audit trail.
Failure DNA captures reusable characteristics of a failure so that future incidents can be compared against previously observed patterns. The audit system provides a tamper-evident history of important incident transitions and actions.
The purpose is not simply to close today's incident. It is to make the system more difficult to fool with the same failure pattern tomorrow.
7. Persists incident state with Firestore
AgentRescue uses Google Cloud Firestore for persistent application state.
The incident record can retain information such as incident identity, taxonomy, verification results, timestamps, and other operational metadata. This gives the control plane a persistent state layer instead of treating every execution as an isolated demo.
8. Runs on Google Cloud
The application is designed and deployed around Google Cloud infrastructure, including Cloud Run and Firestore, with Gemini and Google ADK providing the AI/agent layer.
The repository includes the corresponding Google Cloud deployment configuration and Firestore persistence adapter.
How we built it
AgentRescue was built as a full-stack autonomous-agent security control plane rather than as a single prompt or chatbot.
The major layers are:
UNTRUSTED INPUT
│
▼
┌─────────────────────┐
│ Agent Execution │
│ │
│ Research → Risk → │
│ Procurement → │
│ Purchase Action │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Zero-Trust Gateway │
│ │
│ Identity │
│ Policy │
│ Scope │
│ Tool Evaluation │
└──────────┬──────────┘
│
BLOCK / ALLOW
│
▼
┌─────────────────────┐
│ Incident Control │
│ │
│ Detection │
│ Blast Radius │
│ Investigation │
│ Remediation │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Verification Lab │
│ │
│ Adversarial Tests │
│ Regression Tests │
│ Compliance Tests │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ Human Authorization │
└──────────┬──────────┘
│
▼
RECOVERY
│
┌───────────┴───────────┐
▼ ▼
Failure DNA Audit Trail
│ │
└───────────┬───────────┘
▼
Firestore
AI and agent layer
AgentRescue uses Gemini through the Google GenAI SDK and Google ADK for agent orchestration.
The repository contains an ADK investigation implementation using ADK agents, FunctionTool, Gemini, and InMemoryRunner. The investigation tools are intentionally oriented toward analysis rather than unrestricted execution.
The configured runtime uses Gemini 3.7 Flash for the demonstration environment.
Application layer
The application is implemented as a Node.js / TypeScript web application with an Express backend and Vite-based frontend.
The backend exposes dedicated APIs for:
- agent workflow execution
- ADK investigation
- Gemini incident analysis
- verification
- gateway evaluation
- blast-radius calculation
- Failure DNA generation
- memory search
- incident orchestration
This separation keeps the security controls as explicit application capabilities rather than hiding everything inside one model prompt.
Security architecture
The most important architectural decision was to separate agent reasoning from authorization.
Agents can reason about an action. The gateway decides whether that action is permitted. This prevents an agent from effectively becoming its own security boundary.
The architecture therefore treats the agent fleet as potentially fallible and places policy enforcement outside the model's reasoning loop.
Persistence and memory
Firestore provides persistent storage for application state and incident information. Failure DNA provides a higher-level memory mechanism for recognizing reusable failure patterns. The audit trail records important transitions so that an incident can be reconstructed rather than relying only on the model's narrative.
Challenges we ran into
The hardest part was not making an agent produce an answer. It was designing a system where the agent could be wrong without the entire system becoming unsafe.
Challenge 1: Modeling propagation instead of isolated failure
A prompt injection occurring inside one agent is relatively easy to demonstrate. The harder problem is showing how the contamination affects downstream agents.
We therefore modeled the workflow as a graph and tracked how outputs become inputs for subsequent agents. This allowed us to reason about blast radius instead of only detecting the original malicious text.
Challenge 2: Separating intelligence from authority
It is tempting to give the investigation agent access to everything required to fix the incident. We deliberately avoided that design.
The investigation layer can analyze evidence and recommend remediation, but authorization remains outside the investigator. This required more explicit state transitions and security boundaries, but produced a much stronger architecture.
Challenge 3: Making verification meaningful
A green checkmark is not evidence that a security fix works. We needed the verification layer to test the remediation against multiple classes of failure. That led to the adversarial, regression, and compliance matrix used in the demonstration.
Challenge 4: Making the demonstration honest
AgentRescue demonstrates a security architecture using a controlled scenario. The simulated procurement amount and enterprise action are intentionally not presented as real financial transactions.
This distinction matters because a security product should never manufacture operational credibility by implying that a controlled demo represents a real-world incident. The system is therefore designed to make the boundary between simulation, persisted application state, and real infrastructure explicit.
Challenge 5: Reproducibility and deployment
Integrating a full-stack application with Gemini, ADK, Firestore, and Google Cloud introduced deployment and dependency-management complexity. We documented the deployment architecture and local setup rather than hiding those constraints.
The repository currently uses Bun's lockfile for dependency resolution, while the supplied Cloud Run build configuration has a package-lock/Docker build-contract mismatch that should be resolved before treating it as a guaranteed one-command rebuild. We chose to document this limitation explicitly rather than claim reproducibility that the current snapshot cannot guarantee.
Accomplishments that we're proud of
The accomplishment we are most proud of is that AgentRescue does not stop at detecting prompt injection.
It demonstrates an operational lifecycle:
UNTRUSTED INPUT
↓
PROPAGATION
↓
INCIDENT DETECTION
↓
BLAST-RADIUS ANALYSIS
↓
ZERO-TRUST BLOCK
↓
GOOGLE ADK INVESTIGATION
↓
REMEDIATION
↓
100/100 VERIFICATION
↓
HUMAN AUTHORIZATION
↓
RECOVERY
↓
FAILURE DNA + AUDIT TRAIL
↓
PERSISTENT MEMORY
We are particularly proud of three design decisions.
1. The gateway is independent of agent judgment The system does not assume that an intelligent agent is automatically a trustworthy agent. Authorization is enforced independently.
2. Verification precedes recovery The system does not treat a proposed fix as a verified fix. It creates evidence through a structured test matrix before allowing the workflow to progress.
3. Incidents become institutional knowledge Failure DNA and persistent state mean that an incident can become a reusable security signal rather than disappearing when the workflow ends.
That is the difference between building a demo that handles one attack and designing a control plane that can learn from repeated failure patterns.
What we learned
We learned that building reliable autonomous agents is less about making models more confident and more about designing systems that remain safe when the model is not confident — or is simply wrong.
The most important lesson was:
Model intelligence and system authority should be separate concerns.
A model can be excellent at investigation and still be inappropriate as the final authority for a sensitive action.
We also learned that security for agentic systems cannot be reduced to input filtering. The important questions are:
- Where did the untrusted information enter?
- Which agents consumed it?
- What did it influence?
- What action did it eventually produce?
- Where was that action stopped?
- Was the remediation actually tested?
- Who authorized recovery?
- Will the system recognize the same failure next time?
These questions pushed us toward a control-plane architecture rather than another prompt-defense tool.
We also learned that auditability is part of reliability. If an autonomous system cannot explain what happened, reconstruct the propagation path, and preserve the evidence required to evaluate the response, it becomes extremely difficult to operate responsibly at enterprise scale.
Finally, we learned that being explicit about limitations is itself an engineering feature. AgentRescue is a controlled prototype. It demonstrates the architecture and security lifecycle without pretending that a simulated enterprise transaction is a real-world incident.
What's next for AgentRescue
AgentRescue is currently a controlled prototype demonstrating the architecture. The next stage would be moving from a controlled scenario toward a production-grade agent security platform.
1. Enterprise Agent Registry Introduce a central registry for discovering, versioning, approving, and governing organizational agents. Every agent would have an explicit identity, capabilities, ownership, version, and policy profile.
2. Stronger agent identity and authorization Move from application-level identity metadata toward production-grade workload identity and short-lived credentials. Authorization would become capability-based and scoped to individual tools and resources.
3. Real asynchronous incident operations Extend the current incident lifecycle into long-running background workflows using Google Cloud infrastructure such as Cloud Run, Pub/Sub, and additional managed services. The goal would be to allow AgentRescue to monitor and investigate agent fleets continuously rather than only during an interactive session.
4. Production-grade observability Integrate OpenTelemetry-compatible traces and structured security telemetry to reconstruct complete agent execution chains. This would allow operators to answer not only what an agent did, but why the system allowed it to reach that point.
5. Expanded Failure DNA Failure DNA could evolve into a continuously growing security knowledge base. New incidents could be compared against previous failures, allowing the platform to identify recurring prompt-injection strategies, tool-poisoning patterns, abnormal agent behavior, and policy violations.
6. Broader attack coverage Future verification matrices would expand beyond prompt injection to include:
- tool poisoning
- indirect prompt injection
- malicious RAG content
- privilege escalation
- credential misuse
- cross-agent context contamination
- data exfiltration attempts
- policy bypasses
7. Policy-as-code Organizations could define explicit policies such as:
"Agents may recommend purchases, but cannot authorize purchases above a defined threshold."
or:
"Untrusted retrieved content may influence analysis but may not directly authorize a tool call."
The gateway would enforce those policies consistently across the fleet.
The vision
The long-term vision for AgentRescue is not to make autonomous agents afraid to act.
It is to make them safe enough to act.
As organizations delegate more real work to autonomous systems, reliability cannot depend solely on whether a model gives the right answer. We need systems that can detect when something goes wrong, contain the damage, investigate independently, verify the recovery, require appropriate human authority, and remember what they learned.
--- *An agent can recommend an action. Recommendation is not authorization.*Built With
- adk
- cloudrun
- gemini
- google-cloud
- typescript
Log in or sign up for Devpost to join the conversation.