Inspiration
Most AI governance today exists only as static paperwork, spreadsheets, and manual compliance checklists. Meanwhile, production LLMs and autonomous agents hallucinate critical facts, leak confidential personal data, exhibit demographic bias, invoke unauthorized tools, and remain vulnerable to prompt injection attacks. Compounding the problem, existing evaluation tools lean heavily on subjective "LLM-as-a-judge" prompts that drift over time and leave ephemeral, easily altered logs. These approaches fall short of the deterministic reproducibility required for legal compliance and rigorous regulatory scrutiny.
We built Aegis AI to bridge the divide between theoretical policy and observable runtime behavior. Our guiding philosophy is simple: deterministic where possible, model-assisted where useful, and evidence-backed everywhere. We set out to turn governance into an automated, verifiable software engineering discipline where every model finding is reproducible, auditable, and backed by mathematical certainty.
What it does
Aegis AI is an open-source AI assurance, red-teaming, and governance platform that continuously audits LLMs and autonomous AI agents through tamper-evident cryptographic proof. Instead of treating evaluation as a one-time questionnaire, Aegis observes inference traces, generates targeted probe suites, executes model testing, and evaluates findings across dedicated engines to deliver continuous risk telemetry.
At the core of the platform are six pure assurance engines that operate deterministically. The Fairness engine calculates disparate impact ratios using the EEOC four-fifths rule alongside counterfactual parity tests. The Grounding engine scores RAG claims and citation validity via natural language inference entailment. The Safety engine applies deterministic refusal heuristics and harm classifiers. The Privacy engine scans for sensitive entities using regex and named entity recognition to verify automated redaction. The Security engine subjects models to prompt injection, jailbreaks, and canary token leakage probes. Finally, the Agent Auditing engine monitors multi-agent tool execution to catch unauthorized API calls and boundary violations.
Every evaluation result is committed to a cryptographic evidence vault organized as an append-only SHA-256 hash chain. Database-level triggers permanently block update and delete operations on evidence records, creating tamper-evident audit trails suitable for legal verification. Additionally, an integrated Policy-as-Code compiler ingests regulatory and organizational documents in PDF, Word, or Markdown format and compiles their requirements into executable test controls aligned with NIST AI RMF, ISO/IEC 42001, the EU AI Act, and OWASP LLM Top 10.
Aegis runs completely offline using local models through Ollama with automatic chain-of-thought suppression, ensuring zero cloud API costs and zero data leakage. Teams can instrument production systems in three lines of Python using our client SDK, inspect audit health inside IDEs via our Anthropic Model Context Protocol server, and monitor live evaluations through an interactive 3D neural assurance console.
How we built it
We designed Aegis AI as a modular monorepo that strictly decouples pure mathematical evaluations from web runtimes and database storage. The user interface is built on Next.js 16 and React 19 using Tailwind CSS 4 and Three.js for interactive 3D assurance telemetry. A dedicated Backend-for-Frontend runtime proxy handles session security and streams live audit events over Server-Sent Events directly from our backend workers.
The API backend is written in Python 3.12 using FastAPI to deliver high-throughput asynchronous execution across dozens of endpoints. Long-running audit routines and probe suites are dispatched to solo-pool Celery workers backed by Redis. The evaluation logic lives in an isolated engines package containing pure Python functions with zero database or framework dependencies, ensuring all evaluators remain fully unit-testable and mathematically deterministic.
Data persistence is managed by PostgreSQL 16 with pgvector across 61 schemas. Multi-tenant isolation is enforced at the database engine level through PostgreSQL Row-Level Security policies rather than error-prone application filters. Cryptographic integrity is maintained via procedural database triggers that compute consecutive SHA-256 digests on incoming verdicts and reject any attempts to modify or delete historical records. We also packaged an official Python SDK and an Anthropic Model Context Protocol server to allow external coding agents to trigger compliance runs autonomously.
Challenges we ran into
A major hurdle was eliminating non-determinism during local model evaluations. Small reasoning models like Qwen 2.5 and Qwen 3 running on Ollama frequently output internal reasoning tokens that introduce evaluation drift and latency variance. We resolved this by building automatic chain-of-thought suppression into our local inference adapter and isolating statistical scoring logic from model generations so verdicts remain repeatable across identical inputs.
Implementing court-admissible immutability within a high-concurrency architecture also proved challenging. Relying on application-level soft deletes and permission checks left room for regressions. We solved this by pushing audit chain logic directly into PostgreSQL procedural triggers, ensuring that evidence rows are cryptographically linked and immutable regardless of whether writes originate from Celery workers, API endpoints, or direct database connections.
Finally, orchestrating real-time telemetry across distributed workers required careful backpressure management. We had to stream probe execution updates from Celery workers through Redis pub/sub and FastAPI SSE channels into our 3D Three.js frontend without causing state desynchronization or dropping frame rates during heavy audit runs.
Accomplishments that we're proud of
We succeeded in building an end-to-end assurance pipeline that translates unstructured legal documents directly into executable test suites mapped to international standards like ISO/IEC 42001 and the EU AI Act. This allows compliance teams to verify regulatory posture through software assertions rather than static documentation.
We are equally proud of our local-first implementation. By engineering full support for local Ollama runtimes, Aegis AI delivers enterprise-grade AI red-teaming and compliance testing on standard developer hardware with zero third-party API dependencies. Teams working with sensitive, air-gapped data can run exhaustive safety and privacy audits without transmitting prompts over the public internet.
Lastly, releasing an integrated Model Context Protocol server alongside our core platform allows developers to interact with Aegis directly inside tools like Claude Desktop and Cursor, making AI governance a natural part of daily coding workflows.
What we learned
Translating legal and ethical mandates into software taught us that regulatory compliance becomes manageable when broken down into concrete statistical primitives, such as disparate impact ratios, entailment scores, and tool schema constraints. Abstract policy requirements can be engineered into reliable code when paired with clear operational boundaries.
We also learned that using large language models as judges without deterministic boundaries introduces unacceptable variance. Models are valuable for probe generation and qualitative assistance, but reliable assurance requires combining model outputs with statistical heuristics, rule-based verification, and cryptographic record-keeping.
What's next for Aegis AI
Our immediate priority is shipping pre-built CI/CD release gate plugins for GitHub Actions and GitLab CI, enabling engineering teams to block pull requests and deployment pipelines whenever safety or fairness scores drop below defined thresholds. We are also expanding our policy compiler with ready-to-use compliance packs tailored to sector-specific requirements such as HIPAA, FINRA, and NYC Local Law 144.
Looking further ahead, we plan to expand our agent auditing engine into multi-turn stateful sandboxes capable of stress-testing agent swarms against memory poisoning and cross-agent privilege escalation. We are also investigating zero-knowledge compliance proofs to allow organizations to mathematically certify model safety to external auditors without disclosing underlying prompts, datasets, or proprietary model weights.
Built With
- anthropic
- celery
- docker
- fastapi
- mcp
- next.js
- ollama
- openai
- pgvector
- postgresql
- python
- react
- redis
- tailwind-css
- three.js
- typescript


Log in or sign up for Devpost to join the conversation.