Inspiration

Traditional AI safety and red-teaming frameworks evaluate jailbreaks as syntax flaws or prompt injection bugs. However, frontier model vulnerability in production is increasingly behavioral: multi-turn psychological coercion, manufactured urgency, empathy exploitation, and compliance drift. When interacting across sustained context windows, autonomous agents are prone to progressive boundary decay.

As a clinician and behavioral safety researcher, I designed the AI Analyst-Interceptor framework to treat multi-turn alignment decay with deterministic runtime boundaries rather than non-binding conversational refusals.

What It Does

The AI Analyst-Interceptor Dual Agent Framework is an out-of-band runtime governance architecture that decouples behavioral evaluation from tool execution:

  • The Analyst Agent: Continuously monitors live conversation state, calculates multi-turn semantic drift velocity ($0.00 - 1.00$), and classifies attack vectors across a specialized 31-Law behavioral taxonomy.
  • The Interceptor Agent: Operates as a deterministic, pre-execution circuit breaker. When drift velocity crosses defined thresholds, it intercepts the session, halts state corruption, and forces safe state re-assertion before any backend API, tool call, or data mutation can execute.
  • Cloud Telemetry Pipeline: Automatically streams non-repudiable audit logs to Google Cloud Firestore, providing immutable records of containment, drift scores, and timestamps.

How We Built It

  • Core Governance: Python-based dual-agent execution runtime.
  • Telemetry & State Logging: Google Cloud Firestore (red_team_audits collection) for tamper-evident tracking of adversarial containment.
  • Infrastructure: Google Cloud Platform (Cloud Shell, Cloud Run/Functions compatibility).
  • Intelligence Layer: Gemini API powering multi-turn contextual drift modeling.

Challenges We Faced

A core challenge was eliminating latency bottlenecks while maintaining out-of-band pre-execution governance. Ensuring the Interceptor halts tool execution deterministically—without introducing conversational loops or leaking internal system state—required refining state validation boundaries.

Accomplishments & Empirical Results

In empirical stress testing against sophisticated multi-turn attack vectors (persona hijacks, empathy-bypass exploits, manufactured crisis loops, and historical state gaslighting):

  • Maintained a 100% pre-execution containment rate.
  • Zero downstream tool poisoning or state leakage.
  • Successfully generated real-time, non-repudiable audit records in Cloud Firestore for 100% of tested breach attempts.

What We Learned

Deterministic boundaries must operate out-of-band. When guardrails rely strictly on the primary generative agent's self-governance, semantic drift will eventually bypass conversational alignment. Runtime safety requires decoupled, auditable oversight.

What's Next

  • Expanding the 31-Law behavioral taxonomy into open-source enterprise evaluators.
  • Integrating custom runtime metrics with agent observability and guardrail platforms.
  • Scaling real-time streaming telemetry across high-throughput enterprise agent fleets.

* Providing strategic consulting and coaching for technology companies and healthcare organizations to protect end users from behavioral AI risks while maintaining viable, scalable commercial growth.

Intellectual Property & Legal Notice

  • Patent Status: Patent-Pending (U.S. & International Provisions apply).
  • Copyright Notice: © 2026 Dr. Aunjuli Hicks / Hicks Group Wellness Center. All rights reserved. The 31-Law Threat Taxonomy, Analyst-Interceptor runtime architecture, and dual-agent governance mechanisms are proprietary intellectual property.

Built With

Share this project:

Updates

Submission history