SENTINEL — Evidence-Driven Autonomous Security Verification Fleet

Prove it's broken. Fix it. Prove it's fixed.

Inspiration

Modern software security has an uncomfortable asymmetry: scanners can produce thousands of findings, but a security engineer still has to determine which findings actually matter, reproduce the condition, decide whether a remediation is safe, verify that the remediation really worked, and explain the entire decision months later during an audit.

A finding is therefore not a fact.

It is a claim.

And a remediation is also not a fact.

It is another claim.

SENTINEL was built around one principle:

A security claim should not be trusted merely because an AI model, scanner, or developer says it is true. It should earn its status through evidence.

That idea became the foundation of the system:

[ \boxed{ \text{FIND} \rightarrow \text{UNDERSTAND} \rightarrow \text{VERIFY} \rightarrow \text{REMEDIATE} \rightarrow \text{RE-VERIFY} \rightarrow \text{PROVE} } ]

Instead of building another security chatbot or another vulnerability dashboard, we built an autonomous verification fleet whose job is to turn an uncertain security finding into a machine-observed, reproducible, auditable conclusion.


The Problem

A dependency or application security scan can easily produce dozens or hundreds of alerts.

But a high-severity label does not answer the questions that actually matter:

  • Is the vulnerable component used by this application?
  • Is the vulnerable function reachable?
  • Is the attacker-controlled condition actually present?
  • Can the vulnerability be reproduced against this exact repository state?
  • Will the obvious remediation break existing behavior?
  • After the change, did the vulnerable condition actually disappear?
  • Can the organization prove, months later, why the finding was closed?

SENTINEL targets the gap between:

[ \text{Detection} \neq \text{Verification} \neq \text{Remediation} \neq \text{Proof} ]

The system therefore treats the security lifecycle as an evidence problem rather than a notification problem.


What We Built

SENTINEL is an autonomous security-engineering fleet that investigates a repository, grounds findings against public vulnerability intelligence, reasons about application-specific relevance, executes controlled verification against a real repository checkout, generates a remediation, independently re-runs the original verification scenario, and seals the resulting evidence into a tamper-evident record.

The system is deliberately built as a closed proof loop:

                 ┌───────────────────────────────┐
                 │      SECURITY FINDING         │
                 │ advisory / scanner / manifest │
                 └───────────────┬───────────────┘
                                 │
                                 ▼
                       ┌─────────────────┐
                       │     HUNTER      │
                       │ Find + Ground   │
                       └────────┬────────┘
                                │
                                ▼
                       ┌─────────────────┐
                       │     ANALYST     │
                       │ Is it relevant?│
                       └────────┬────────┘
                                │
                                ▼
                  ┌─────────────────────────────┐
                  │      VERIFICATION LAB       │
                  │ isolated real-code testing  │
                  └─────────────┬───────────────┘
                                │
                         CONFIRMED / REJECTED
                                │
                                ▼
                       ┌─────────────────┐
                       │   PATCH FORGE   │
                       │   remediation   │
                       └────────┬────────┘
                                │
                                ▼
                       ┌─────────────────┐
                       │  RE-VERIFIER    │
                       │ same scenario   │
                       │ after the fix   │
                       └────────┬────────┘
                                │
                                ▼
                       ┌─────────────────┐
                       │ EVIDENCE AGENT  │
                       │ seal + archive  │
                       └────────┬────────┘
                                │
                                ▼
                    ┌──────────────────────────┐
                    │ HUMAN DEPLOYMENT GATE    │
                    │ approve / reject         │
                    └──────────────────────────┘
The core implementation already exists in the repository as six specialized agents: Hunter, Analyst, Verification Lab, Patch Forge, Re-Verifier, and Evidence Agent.

Why This Is Different

The individual technologies are not the novelty.

AI code generation already exists.

Security scanning already exists.

Dependency databases already exist.

Sandboxed testing already exists.

PDF processing already exists.

Digital signatures already exist.

The differentiation is the composition of these capabilities into an evidence-producing control loop.

SENTINEL does not allow a model's narration to become a security verdict.

The Analyst may propose that a finding is relevant, but the system does not treat that proposal as proof. The Verification Lab must execute a real verification scenario and return an observed result before the finding can be considered confirmed. The same principle is applied after remediation.

This creates an important invariant:

$$ \boxed{ \text{AI Opinion} \not\Rightarrow \text{Verified Security Fact} } $$

Instead:

$$ \boxed{ \text{Verified Fact} = f( \text{source evidence}, \text{code observation}, \text{controlled execution}, \text{post-fix re-execution} ) } $$
1. HUNTER — FINDINGS WITH GROUNDING

Hunter discovers security findings from repository state and vulnerability intelligence.

SENTINEL grounds advisory identifiers against live security knowledge sources including:

OSV
NVD
GitHub Advisory Database
EPSS

An advisory that cannot be resolved is not silently treated as a clean result. The system distinguishes an actually unresolved advisory from a degraded external lookup condition.

This prevents a dangerous failure mode:

$$ \text{External failure} \not\equiv \text{No vulnerability} $$

That principle is important because an autonomous security system should fail loudly, not produce a reassuring but incorrect "zero findings" result.

2. ANALYST — APPLICATION-SPECIFIC RELEVANCE

A CVE or advisory can be real while still being irrelevant to a particular application.

SENTINEL therefore examines application context instead of blindly trusting severity labels.

The Analyst combines:

Finding
   +
Source code
   +
Dependency relationship
   +
Reachability
   +
Configuration
   +
Grounded advisory information

The result is a structured verdict with claims and supporting sources.

For example:

Finding:
GHSA-8cf7-32gw-wr33

Claim:
jsonwebtoken is a direct dependency.

Evidence:
trace_reachability:lib/insecurity.ts:11

Claim:
the advisory describes the vulnerable behavior.

Evidence:
osv:GHSA-8cf7-32gw-wr33

Claim:
jwt.verify is invoked without restricting algorithms.

Evidence:
verified_fixes:CWE-327

The repository implements this source-aware evidence model explicitly.

3. VERIFICATION LAB — THE SYSTEM DOES NOT ASK YOU TO TRUST THE MODEL

This is the heart of SENTINEL.

The Verification Lab does not generate a sentence such as:

"This vulnerability appears exploitable."

It executes a controlled test against a genuine checkout of the repository.

The verification scenario records:

{
  "finding_id": "F-1042",
  "scenario": "reachability:/api/checkout:sql_injection",
  "expected": "vulnerable_function() should be reachable",
  "observed": "vulnerable_function() reached",
  "result": "CONFIRMED_EXPLOITABLE",
  "sandbox_id": "runsc-8841",
  "duration_ms": 4210
}

The repository explicitly uses real repository state and real command execution rather than simulated results.

The design principle is:

$$ \boxed{ \text{Verdict is earned by execution} } $$

not:

$$ \boxed{ \text{Verdict is generated by narration} } $$
4. PATCH FORGE — REMEDIATION WITH A BOUNDARY

Once a condition has been verified, Patch Forge generates a remediation.

But SENTINEL deliberately does not give this agent unrestricted production authority.

Its permission boundary is closer to:

READ repository
+
READ security knowledge
+
CREATE remediation branch
-
NO production deployment

Fixes are based on catalogued OWASP-derived patterns. If there is no supported remediation pattern for a weakness class, SENTINEL escalates rather than improvising a potentially unsafe change.

This creates another safety property:

$$ \text{Unsupported remediation} \rightarrow \text{Human Review} $$

instead of:

$$ \text{Unsupported remediation} \rightarrow \text{LLM guess} $$
5. RE-VERIFIER — PROVING THE FIX ACTUALLY WORKED

A version bump is not proof.

A modified line of code is not proof.

A passing LLM explanation is not proof.

SENTINEL takes the original verification scenario and re-runs it after the remediation.

Conceptually:

$$ V_{\text{before}} = \text{Execute}(C_{\text{original}}, S) $$ $$ V_{\text{after}} = \text{Execute}(C_{\text{patched}}, S) $$

The remediation is only considered successful when the observed condition changes from the vulnerable state to the resolved state and regression checks remain successful.

Therefore:

$$ \boxed{ \text{Verified Fix} = V_{\text{before}}=\text{FAIL/EXPLOITABLE} \land V_{\text{after}}=\text{RESOLVED} \land R=\text{PASS} } $$

This is the core "prove it twice" philosophy of SENTINEL.

6. EVIDENCE AGENT — FROM AI OUTPUT TO PROVABLE ARTIFACT

Every important action becomes part of an evidence chain.

Finding
   ↓
Grounded verdict
   ↓
Verification result
   ↓
Patch proposal
   ↓
Re-verification result
   ↓
EvidenceObject
   ↓
SHA-256 signature
   ↓
Signed PDF
   ↓
Archive

The repository's evidence pipeline uses both a content signature and a DWS document signature. The JSON evidence record is hashed independently, while the rendered PDF receives its own digital signature.

The result is a dual-integrity model:

$$ H_{\text{record}} = SHA256(EvidenceObject) $$ $$ H_{\text{document}} = SHA256(PDF) $$

The two artifacts are independently verifiable.

This matters because a system should not merely say:

"Here is the report."

It should make it possible to ask:

"Has the record or the document changed since it was sealed?"

A Proof-Carrying Security Record

The final output is therefore more than an alert.

It is closer to a proof package:

┌─────────────────────────────────────────────────┐
│           SENTINEL SECURITY EVIDENCE             │
├─────────────────────────────────────────────────┤
│ Finding            GHSA-...                     │
│ Repository         payment-service              │
│ Commit             a81f92c                      │
│                                                   │
│ Grounding          ✓                            │
│ Reachability       ✓                            │
│ Verification       ✓                            │
│ Remediation        ✓                            │
│ Re-verification    ✓                            │
│ Regression         ✓                            │
│ Evidence seal      ✓                            │
│ Human decision     PENDING                      │
└─────────────────────────────────────────────────┘

The evidence model and replay timeline are implemented as part of the repository rather than merely described as future functionality.

7. GOVERNANCE — THE AGENTS ARE NOT TRUSTED BY DEFAULT

One of the most important engineering choices in SENTINEL is that the agent fleet is governed through explicit policy.

Every tool call passes through an Agent Gateway.

                AGENT
                  │
                  ▼
        ┌──────────────────┐
        │  AGENT GATEWAY   │
        └────────┬─────────┘
                 │
        ┌────────┼────────┐
        ▼        ▼        ▼
     Registry Identity Model Armor
        │        │        │
        └────────┼────────┘
                 ▼
          Tool Execution
                 │
                 ▼
            Audit Log

The implementation checks agent registration status, identity scope, and model-safety policies. An unapproved agent can be denied even when its technical identity would otherwise allow the action.

A powerful example is the Deployment Gate:

patch-forge → deploy to production
                     │
                     ▼
              REQUIRES_HUMAN

No agent owns the final production permission.

This creates a deliberate trust boundary:

$$ \boxed{ \text{Autonomy where computation is repeatable} \quad+\quad \text{Human authority where consequences are irreversible} } $$
8. ASYNCHRONOUS AGENT RUNTIME

Security investigations are not always short API calls.

A scan can take minutes.

A repository can be cloned after the browser has already closed.

A worker can fail halfway through.

SENTINEL therefore models an investigation as a durable asynchronous job:

POST investigation
       ↓
   Job Queue
       ↓
     Worker
       ↓
 Agent Fleet
       ↓
 Evidence Sealed
       ↓
 Human Review

The API enqueues work instead of performing the entire investigation synchronously, allowing the browser to disconnect without destroying the job. The repository also implements job leasing and recovery from abandoned work.

This transforms the system from:

"an AI request"

into:

"a long-running autonomous engineering operation."

9. MEMORY THAT IMPROVES WITH VERIFIED HISTORY

SENTINEL maintains a memory bank containing prior verdicts, remediation patterns, and verified fixes.

That means a future investigation does not always have to start from a blank state.

Conceptually:

$$ P(\text{best remediation}\mid C) \rightarrow P(\text{best remediation}\mid C,\mathcal{M}_{verified}) $$

where:

\(C\) = current repository context
\(\mathcal{M}_{verified}\) = previously verified remediation memory

The important distinction is that memory is used as grounding, not as unquestioned truth. Previous patterns are still subject to current verification.

The repository implements this using ChromaDB with embedded embeddings and persistent prior investigation knowledge.

10. FAILURE IS A FIRST-CLASS STATE

A major lesson while building SENTINEL was that autonomous systems are dangerous when failures look like successes.

We deliberately designed for failure transparency.

Examples discovered and fixed during development included:

scans hanging when repositories had no lockfile
slow environments exceeding initial timeouts
queue jobs becoming permanently stuck after worker death
cloud clients repeatedly rebuilding expensive connections
model-driven orchestrators finishing without actually sealing evidence
PII detections being visually confused with clean results
static deployments generating invalid prefetch requests

These are not merely theoretical failure modes; the repository documents them as real development incidents and ties them to concrete fixes.

The design invariant is:

$$ \boxed{ \text{Unknown / degraded / failed} \neq \text{safe} } $$

In security engineering, uncertainty must remain visible.

11. A MULTI-INTEGRATION ARCHITECTURE

SENTINEL was designed so orchestration, state, and infrastructure remain interchangeable.

The core investigation logic can run through:

Direct Python
Google ADK
AWS Strands

while using interchangeable backends such as:

Local
Firestore
DynamoDB

Local Queue
Pub/Sub
EventBridge

The same investigation functions are reused instead of maintaining separate security logic for each cloud platform.

This means:

$$ \text{One Security Engine} + \text{Multiple Execution Strategies} $$

rather than:

$$ \text{Google Implementation} \neq \text{AWS Implementation} $$

The current repository explicitly implements this abstraction through factories and orchestration adapters.

12. NUTRIENT DWS — WHY DOCUMENTS ARE PART OF THE SECURITY SYSTEM

Security decisions are rarely consumed only by developers.

They eventually reach:

security leadership
compliance teams
auditors
risk teams
customers
governance committees

SENTINEL therefore turns machine observations into a structured evidence artifact.

Nutrient DWS performs meaningful document operations in the project:

Security Evidence
       ↓
DWS build
       ↓
Auditable PDF
       ↓
DWS signing
       ↓
Tamper-evident artifact

The current repository uses DWS for report rendering and CAdES signing.

This directly aligns with the DevNetwork Nutrient challenge, whose official requirement is meaningful use of DWS for a core document operation, with bonus emphasis on deterministic, auditable output and human review where judgment is required.

13. THE HUMAN IS NOT REMOVED — HUMAN JUDGMENT IS MOVED TO THE RIGHT PLACE

SENTINEL's goal is not:

"Remove humans from security."

It is:

"Remove humans from repetitive investigation while preserving human authority over consequential decisions."

Therefore:

AI:
  discover
  ground
  reason
  execute tests
  generate remediation
  re-test
  package evidence

Human:
  review exceptional cases
  approve/reject deployment

This creates the control principle:

$$ \boxed{ \text{Autonomous Computation} + \text{Human Governance} = \text{Responsible Agentic Security} } $$
14. THE DEVNETWORK API + CLOUD + AI CONNECTION

The DevNetwork hackathon evaluates projects on three dimensions:

PROGRESS

How much was actually built?

CONCEPT

Does it solve a real problem?

FEASIBILITY

Could it become a startup or company?

SENTINEL was engineered around those three questions.

Progress

This is not a mock UI.

The repository contains:

a live dashboard
live backend API
asynchronous job execution
a six-stage investigation pipeline
persistent evidence
governance controls
document sealing
real repository scanning
automated tests

The README provides a live dashboard and live backend and states that starting an investigation executes the real pipeline.

Concept

The problem is universal across software organizations:

$$ \text{Too many findings} + \text{Too little engineering time} + \text{Weak evidence} $$

SENTINEL converts that into an automated evidence workflow.

Feasibility

The buyers are straightforward:

security teams
DevSecOps teams
regulated software organizations
SaaS companies
financial services
enterprise engineering teams

The commercial unit can be:

$$ \text{Repository} \quad\text{or}\quad \text{Verified Finding} \quad\text{or}\quad \text{Assurance Run} $$

That creates a plausible SaaS model without requiring the customer to replace every existing scanner.

SENTINEL can instead become the verification and assurance layer above existing security tooling.

15. THE PRODUCT VISION

The long-term vision is larger than vulnerability remediation.

SENTINEL is evolving toward a broader category:

Evidence-Driven Software Assurance

The future version can reason about:

Security
   │
   ├── vulnerabilities
   ├── secrets
   ├── authentication
   └── authorization

Reliability
   │
   ├── regression
   ├── dependency changes
   └── deployment risk

Compliance
   │
   ├── policies
   ├── controls
   └── audit evidence

AI-Generated Changes
   │
   ├── code generated by agents
   ├── generated configuration
   └── generated infrastructure

The deeper idea is:

When software becomes increasingly generated by autonomous systems, we need an equally autonomous system whose job is to independently prove that the generated change is safe.

That is the strategic direction behind SENTINEL.

16. WHAT WE LEARNED BUILDING IT

The hardest lessons were not about prompting.

They were about systems engineering.

We learned that an autonomous agent system needs:

explicit state transitions
immutable evidence
least-privilege permissions
failure-aware queues
reproducible execution
grounded knowledge
independent verification
human approval boundaries
observable agent behavior
strong contracts between agents

A model can produce a convincing answer in milliseconds.

A production system must also answer:

What happened?

What evidence supports it?

Which tool executed?

Which repository state was tested?

What changed?

Did the fix really work?

Who approved it?

SENTINEL was built to answer those questions.

17. WHY "SHOW YOUR WORK" IS THE PRODUCT

The most important architectural rule in SENTINEL is simple:

$$ \boxed{ \text{No Evidence} \Rightarrow \text{No Trusted Verdict} } $$

The system therefore treats evidence as a first-class data product.

Every important conclusion should be replayable.

Every important action should be attributable.

Every important document should be tamper-evident.

Every high-impact action should have an explicit authority boundary.

This is why SENTINEL is not simply an AI security agent.

It is an evidence-producing autonomous control system.

18. FINAL DEMO

Our end-to-end demonstration is intentionally simple:

1. A real vulnerable repository is submitted.
                    ↓
2. SENTINEL identifies and grounds the finding.
                    ↓
3. Analyst determines application-specific relevance.
                    ↓
4. Verification Lab reproduces the vulnerable condition.
                    ↓
5. Patch Forge creates a remediation branch.
                    ↓
6. Re-Verifier runs the original scenario again.
                    ↓
7. The vulnerability is shown as resolved.
                    ↓
8. Regression evidence is recorded.
                    ↓
9. Nutrient DWS generates and signs the evidence artifact.
                    ↓
10. Human Deployment Gate approves or rejects the result.

The interface is designed to make that sequence visible rather than hiding everything behind a chat window. The current product includes a 3D fleet visualization, Command Center, Verification Lab, Remediation Forge, Evidence Report, Governance surface, Audit Ledger, and Deployment Gate.

19. THE CENTRAL IDEA

Security teams already have scanners.

Developers already have coding agents.

Organizations already have vulnerability databases.

What is missing is a reliable bridge between:

$$ \text{AI-generated conclusion} $$

and

$$ \text{Evidence-backed engineering decision} $$

SENTINEL is that bridge.

We do not ask the organization to blindly trust the agent.

We ask the agent to produce evidence.

SENTINEL — Prove it's broken. Fix it. Prove it fixed.
Let:

F = raw security finding
G = grounded advisory evidence
R = relevance analysis
V = controlled verification
P = remediation
V' = independent post-fix verification
A = audit/evidence integrity

Then:

Trusted Security Decision
= F ∧ G ∧ R ∧ V ∧ P ∧ V' ∧ A

where:

V = execution before remediation
V' = independent execution after remediation

and:

Trusted Fix
⇔
(V = vulnerable)
∧
(V' = resolved)
∧
(regression tests = pass)

Built With

  • a2a-protocol
  • agent-gateway
  • agent-identity
  • agent-registry
  • agent-runtime
  • artifact-registry
  • cloud-run
  • cloud-storage
  • cloud-trace
  • fastapi
  • firestore
  • gemini-3.5+
  • gemini-enterprise-agent-platform
  • google-adk
  • iam
  • mcp
  • memory-bank
  • model-armor
  • next.js
  • nutrientdws
  • opentelemetry
  • pub/sub
  • python-3.12
  • typescript
  • workload-identity-federation
Share this project:

Updates

Submission history