Inspiration

AI systems are moving from generating answers to taking actions.

They can call tools, invoke APIs, modify records, trigger workflows, interact with infrastructure, and reach downstream services. But most assurance systems still focus on individual components: the model, the permission, the API, the identity, or the scanner finding.

That creates a fundamental gap:

What an AI system is supposed to access may be different from what it can actually reach.

A recent incident reported by OpenAI illustrates this class of problem. Agents without direct internet access found an indirect route through Artifactory, used it as an unintended coordination channel, reached the internet, and later reached Hugging Face systems.

The lesson is not that every system contains the same exploit.

The lesson is that declared boundaries and reachable boundaries can be different.

A permission list tells us what a system is supposed to access.

HAIEC asks what the AI-enabled system can actually reach through the code, tools, services, identities, and resources connected to it.

The direct door may be locked while another hallway still exists.

Reference: https://openai.com/index/hugging-face-incident-and-the-road-ahead/


What it does

AI Action Path Assurance reconstructs the evidence-backed path from an AI-facing capability to a real system consequence.

A simplified path looks like:

AI-facing capability
        ↓
Model exposure / registration
        ↓
Dispatch
        ↓
Implementation
        ↓
Handler
        ↓
Resource or service
        ↓
Consequential operation

Where qualifying evidence exists, HAIEC continues further:

Effective arguments
        ↓
Authenticated / tenant context
        ↓
Confirmation or approval mediation
        ↓
Downstream service reachability
        ↓
Persistence and system effects

Instead of collapsing everything into one risk score, HAIEC keeps five assurance questions separate:

Assurance Plane Question
Requested Was this action actually declared or requested?
Policy Authorized Was it approved for this use and scope?
Effectively Granted Did the acting identity actually have the authority?
Code Capable Can the evaluated implementation reach the consequence?
Observed Do qualifying logs, traces, or witnesses show it actually happened?

Capability is not permission. Permission is not execution.

This separation matters because proving one plane does not prove the others.


How we built it

HAIEC is designed around a simple principle:

The assurance claim should come from evidence, not from one AI model judging another.

Not AI testing AI

HAIEC does not ask an LLM, README, policy document, or tool inventory to guess what a system can do.

Under the hood, HAIEC constructs a deterministic evidence graph connecting:

AI capability
→ exposure
→ dispatch
→ implementation
→ resource
→ consequential operation

It then reconciles those paths against independent authority, control, and runtime-evidence planes.

LLMs may help explain the evidence.

They do not establish the assurance claim.

Source-bound analysis

Every evaluation is bound to an exact source snapshot.

Supported evidence can establish relationships such as:

  • AI-facing capabilities
  • model exposure and registration
  • dispatch relationships
  • implementation and handler bindings
  • consequential operations
  • tenant and execution context
  • resource and service reachability
  • evidence coverage
  • artifact integrity

Enterprise input is incorporated where only the organization can define proprietary business meaning, approved operating boundaries, policy, effective authority, or runtime evidence.

Verified Action Frontier

Not every assurance question can always be answered from source code alone.

Instead of hiding that uncertainty inside a confidence score, HAIEC records exactly how far each action path has been verified.

We call this the Verified Action Frontier.

It shows:

  • what is established,
  • what remains unresolved,
  • what evidence is missing,
  • and what evidence would strengthen the claim next.

This turns uncertainty into actionable assurance work.


Real-world validation: KestrelVoice

We tested the system against KestrelVoice, a working multi-tenant AI SaaS application where AI agents use tools and application handlers to perform consequential application actions.

Frozen source commit:

5e65843fddfe5f907485b798e464ad37b3b3b2c7

Public evaluation:

HAIEC-KESTREL-EVAL-5e65843

Assurance disposition:

REVIEW

Frozen evaluation result

  • 44 consequential action paths confirmed in code
  • 44 / 44 evaluated consequential paths satisfied the Code Capable evidence requirement
  • 1,322 supporting Action Path evidence traces
  • 555 Verified Action Frontier items
  • Requested: NOT_ASSESSED
  • Policy Authorized: NOT_ASSESSED
  • Effectively Granted: NOT_ASSESSED
  • Observed: NOT_ASSESSED

One representative path:

AI service-call agent
        ↓
AI-facing capability
        ↓
handle_place_order
        ↓
orders table insert
        ↓
ORDER RECORD WRITE

HAIEC established that this source-backed path exists.

It does not automatically claim that the action was organizationally authorized, permitted at runtime, executed in production, or completed as a customer transaction.

Those remain separate evidence questions.


Path-specific control evidence

A system-wide statement such as:

“We have tenant isolation.”

is not sufficient for consequential-action assurance.

HAIEC asks a narrower question:

Is the control established on the exact path that reaches the consequence?

In the evaluated Kestrel evidence, tenant filtering was verified on 30 of 31 applicable relations.

For the ORDER RECORD WRITE consequence relation, tenant context was present, but tenant filtering and authenticated-subject binding were not established on that specific relation.

Instead of concluding that the entire application is safe or unsafe, HAIEC produces a precise evidence gap that can be investigated.


Release A / Release B comparison

Traditional Git diff answers:

What changed in the files?

HAIEC asks:

What changed in consequential reach, controls, and assurance evidence?

We call this Consequence Delta.

For the qualified Kestrel Release A / Release B comparison:

  • Overall result: INCONCLUSIVE
  • 12 CONTROL_CHANGED items established
  • Release perimeter: SAME
  • Analyzer comparability: EXACT
  • Repeatability: PASS
  • 99 positional facts remained UNRESOLVED

The strongest established transition was control evidence moving from:

DECLARED

to:

DECLARED + PATH_BOUND

Where evidence could not prove that a mutation had been added or removed, HAIEC did not manufacture a conclusion.

UNRESOLVED ≠ ABSENT

That distinction is intentional.


Challenges we ran into

1. Separating capability from authority

One of the hardest design problems was preventing evidence from one assurance plane from silently becoming evidence for another.

Source code can establish that an implementation can reach a database write.

It cannot automatically establish that the organization approved the action or that it executed in production.

Keeping these planes independent required strict evidence contracts.

2. Proving absence

Establishing that a path exists is often easier than proving that no alternative path exists.

For release comparison, we therefore distinguish between:

  • established change,
  • unchanged evidence,
  • and unresolved evidence.

We intentionally prefer INCONCLUSIVE over an unsupported conclusion.

3. Following real application paths

AI capabilities rarely map directly to consequences.

The path may cross model registrations, tool routers, handlers, service layers, database abstractions, authentication context, tenant controls, and downstream systems.

Reconstructing those relationships deterministically required building a graph rather than relying on isolated scanner findings.

4. Making technical evidence usable

Developers, executives, auditors, governance teams, and automated systems need different views of the same evidence.

We designed HAIEC so one canonical evidence base can project into multiple outputs without re-analyzing the application separately for every audience.


Accomplishments that we're proud of

We moved beyond a theoretical framework and demonstrated the approach against a real multi-tenant AI application.

Key accomplishments include:

  • Reconstructing 44 consequential AI action paths
  • Establishing Code Capable evidence for 44 / 44 evaluated consequential paths
  • Producing 1,322 source-backed evidence traces
  • Tracking 555 Verified Action Frontier items
  • Building path-specific control analysis
  • Separating five independent assurance planes
  • Building qualified Release A / Release B consequence comparison
  • Binding evidence to exact source snapshots
  • Producing machine-readable and human-readable assurance outputs
  • Refusing unsupported conclusions when evidence is incomplete

The last point is especially important.

For high-consequence AI systems, knowing what the evidence cannot yet establish is itself valuable information.


What we learned

The action path is the real unit of assurance

Looking only at a model, permission, API, or scanner finding misses how consequential actions emerge across systems.

The meaningful object is:

the path from AI exposure to consequence, together with the evidence that bounds what can be claimed about it.

Permission is not delegation

A permission may exist without an AI system being authorized to exercise it for a particular purpose.

Likewise, an AI capability may be technically reachable without being policy-authorized or actually executed.

These relationships must remain separate.

Controls need path-level evidence

A control existing somewhere in an architecture is different from establishing that it mediates the exact path reaching a consequential operation.

Uncertainty should be explicit

Assurance systems often create pressure to output a score, pass/fail label, or confident conclusion.

For consequential AI, sometimes the correct answer is:

NOT_ASSESSED
UNRESOLVED
INCONCLUSIVE

Those states should trigger additional evidence collection rather than disappear inside a confidence score.


What makes AI Action Path Assurance technically different

1. Consequence-centered assurance

The primary unit is the qualified path from an AI-facing capability to a consequential operation.

2. Independent evidence planes

Requested, Policy Authorized, Effectively Granted, Code Capable, and Observed remain independent rather than collapsing into one risk score.

3. Verified Action Frontier

HAIEC records how far each path is verified today and identifies what evidence strengthens the assurance claim next.

4. Consequence Delta

Between qualified evaluations, HAIEC compares changes in consequential reach, controls, and evidence coverage rather than only source-code changes.

5. Portable evidence

One canonical evidence base can project into executive, technical, auditor, and machine-readable views.

6. Evidence integrity

Artifacts are bound to the evaluated source and underlying evidence through manifests and cryptographic digests.


Outputs

HAIEC can produce:

  • Executive Assurance Report
  • Technical Assurance Report
  • Auditor / Evidence Report
  • System Consequence Map
  • Machine-readable evidence
  • Decision Receipt
  • Agentic Production Passport
  • Consequence Delta
  • Evidence manifests and integrity artifacts

Enterprise extensibility

HAIEC starts with a standard capability model for common enterprise operations.

Organizations can extend the model when proprietary operations carry specific business consequences.

Examples:

Telecom
restartCell()
changeAntennaTilt()
publishRAppPolicy()
Banking
releaseWireTransfer()
freezeAccount()
adjustCreditLimit()
Manufacturing
restartProductionCell()
changeRobotSpeed()
publishPLCConfiguration()

These capabilities can then be compared against an approved Operating Envelope describing allowed:

  • operations
  • resources
  • data classes
  • destinations
  • models and tools
  • environments
  • quantitative bounds
  • required approvals

What's next for AI Action Path Assurance

The next stage is moving from source-backed capability assurance toward a continuously maintained AI action assurance layer.

Our roadmap includes:

  1. Runtime evidence correlation Connect qualified source paths with logs, traces, approvals, and other runtime witnesses.

  2. Operating Envelope enforcement Compare AI-reachable actions against organization-defined operational boundaries.

  3. Release assurance Automatically identify material changes to consequential reach or control coverage before production release.

  4. Enterprise capability libraries Expand consequence models for regulated and high-impact domains such as telecom, banking, healthcare, infrastructure, and manufacturing.

  5. Agentic Production Passports Package source identity, consequence reachability, controls, evidence coverage, release changes, and unresolved assurance questions into portable decision artifacts.

  6. Continuous assurance Shift AI governance from periodic documentation toward evidence that changes alongside the system.

The long-term goal is straightforward:

Before an organization delegates consequential action to AI, it should be able to establish what that AI can actually cause, under what authority, through which path, and with what evidence.


Why it matters

Agentic AI changes the unit of assurance.

The important object is no longer only:

  • the model,
  • the tool,
  • the identity,
  • the permission,
  • the scanner finding,
  • or the runtime event.

It is the action path to consequence and the evidence surrounding that path.

HAIEC helps release, security, platform, governance, and assurance teams answer:

  • What can this AI-enabled system actually reach?
  • What material action can it cause?
  • What controls and authority are established on that path?
  • What has actually been observed?
  • What changed in the next release?
  • What remains unresolved?
  • What evidence should travel with the decision?

Public demo and evidence

Kestrel Assurance Report https://www.haiec.com/sample-reports/kestrel

System Consequence Map https://www.haiec.com/sample-reports/kestrel/constellation

Release A / Release B Comparison https://www.haiec.com/sample-reports/kestrel/compare

Share this project:

Updates

Submission history