Inspiration
Every tool I could find for securing AI agents asks the same question: is this agent allowed to do this? It is a good question. It is also a question about one actor at a time.
Then I read what the field was actually reporting. OWASP's Top 10 for Agentic Applications catalogues insecure inter-agent communication. The Cloud Security Alliance published a note on coordination emerging among evaluation agents that were never instructed to cooperate, and another on autonomous agents as a third dimension of lateral movement. None of those describe an agent breaking its rules. They describe a population arriving somewhere no member was sent.
So I wrote down the smallest version of it I could think of. Four lines of log, each one an action its actor is explicitly permitted to take:
Support Agent read Customer CRM ← allowed: support reads customer records
Support Agent write Shared Scratchpad ← a scratchpad is just a scratchpad
Reporting Agent read Shared Scratchpad ← reporting reads what it is given
Reporting Agent send External Webhook ← allowed: reporting is *for* sending reports
No permission was violated. No agent misbehaved. Nothing there would trip a per-action policy engine, a prompt firewall, or an audit of any individual agent's scopes. And restricted customer data can now leave the estate.
The capability belongs to the population, not to any member of it. It was assembled from four permissions that were each correctly granted, and it appeared the moment an ordinary read connected two halves that were never meant to touch.
That is the class of risk HIVE exists to find.
What it does
HIVE takes an architecture manifest — a YAML file declaring the zones, nodes, relationships and invariants an agent estate is supposed to have — and an append-only observation ledger of what the estate actually did. It then does three things, in order, and refuses to skip any of them.
Detect. It builds a behaviour graph from the ledger and runs rules that look for compositions the declared architecture forbids. The headline rule finds a route from restricted data to an external destination that passes through shared state absent from the manifest. Crucially, it gates on manifest deviation: a route made entirely of declared relationships is the architecture working, not a finding.
Contain. It clones the graph, applies each pre-authorised control to the clone, re-runs the detector, and compares the resulting declared-relationship set against the original. Controls are ranked by cost, and any control that removes the path by severing work the architecture declares is rejected and shown as rejected — with the relationship it would have broken named. On the canonical scenario two of the five candidates are refused for exactly that reason.
Verify. After a control is issued, detection runs again over the contained graph. The planner predicting success is not accepted as evidence that it worked.
It also keeps an immunity memory — the abstract structure behind a finding, recorded as a draft that a human must promote — and a Swarm Lab that builds a synthetic population from parameters you choose and runs the same detection and containment code over it, so the answer is produced rather than estimated.
The emergence score is deliberately not a model. It is a weighted sum of indicator factors:
$$E \;=\; \sum_{i=1}^{n} w_i \cdot \mathbb{1}[f_i], \qquad \sum_{i=1}^{n} w_i = 1$$
Weights are fixed and visible, so a reader can reconstruct the number by hand from the factor list. It ranks findings; the factors are what explain them.
There is no LLM anywhere in the runtime. No model call, no API key, no network dependency. The analysis is deterministic Python and behaves identically offline.
How I built it
The core is Python 3.12 with FastAPI, Pydantic v2 and NetworkX. Two design choices did most of the work:
The ledger is the truth; the graph is a derived view. Observations are appended and never mutated, and the graph is rebuilt by replaying them. Rewinding the timeline is not a cached picture — it is a genuine replay, which is what makes a claim about state reproducible rather than asserted.
Controls are evidence. A control does not quietly mutate a graph. It appends a block event to the same ledger, so the record of why the graph changed sits beside the record of what changed, and a replay reconstructs the contained state exactly.
The console is React 19, TypeScript, Vite, Zustand and Tailwind v4. It talks to a live core service when one is running and falls back to a recorded snapshot — the real engine's output captured at every replay position — when one is not. That is how the hosted demo is fully interactive with no backend. It is not a mock: every byte came out of the real engine, and the console says recorded session on screen rather than implying a live service.
CI runs the formatter, the linter, mypy --strict and all 231 tests, plus two drift gates: the committed OpenAPI document is regenerated from the running service, and the recorded demo is regenerated from the engine — both compared byte for byte. If the engine changes and the recording does not, the build fails. The page you can click is the code you can read.
Challenges I ran into
A corrupt fixture that failed silently. An integration test insisted the detector found nothing. The detector was fine. The scenario file had been concatenated from two versions and contained duplicate event IDs — and the ledger, correctly, de-duplicates by ID. Half the baseline was being dropped without a word, so the agent the rule needed had never existed. The fix was one clean fixture, and a loader that now rejects duplicate IDs and duplicate sequence numbers instead of quietly discarding them.
A graph that ate its own edges. I stored the behaviour graph as a DiGraph, so a discover and a write between the same two nodes overwrote each other — one action silently replaced the other. It became a MultiDiGraph keyed by action, and containment became scoped to the action rather than to the pair of nodes.
Simulation that did not match reality. The planner simulated containment action-agnostically but applied it action-scoped. It could predict one outcome and produce another — which would have quietly invalidated the entire containment argument. Now simulation and application call the same apply_control function, and a test applies all five registered controls for real and compares each against what was predicted.
Data that travels against the arrow. An invariant kept reporting a violation after containment had demonstrably removed the path. The graph records Reporting Agent --read--> Scratchpad, but the data moves the other way, and discover carries no data at all. Once reads were re-oriented and discovery excluded from the data-flow view, the rule and the invariant agreed at all twelve replay positions.
A rule that was unreachable, then a rule that cried wolf. The second detector matched on an action the schema validator rejected, so it was dead code that had never once fired. Fixing that made it fire on the declared baseline — correct behaviour flagged as a finding. It needed the manifest-deviation gate: at least one hop in the composition must be something the architecture does not declare.
A demo that would have failed my own gate. I added a CI check that the recorded demo must be byte-identical when regenerated. It immediately failed — three times over. A generation timestamp, Prettier reformatting the generated JSON, and, worst, datetime.now() stamped into control and immunity records. A ledger whose contents depend on when someone clicked cannot be replayed byte-for-byte. Controls are now stamped from the replay position, and a test asserts the whole ledger — the control HIVE issued and the pattern it recorded included — comes out identical across two detect-contain cycles.
Browser tests that fought each other. The service holds replay state per scenario, not per client, so parallel end-to-end tests interfered. I set the worker count to one and documented why in the config, rather than papering over a real property of the service.
Accomplishments I'm proud of
- 231 tests, organised around what HIVE claims rather than around its modules — including negative tests that remove each precondition in turn and require the rule to go silent, and property-based tests over arbitrary event streams.
- A byte-identical ledger across replays, which is what lets me call the demo evidence instead of a recording.
- Containment that is allowed to say no. The two controls that would sever declared work are surfaced and refused. Hiding them would have been easier and worth less.
- Findings that state their own limits. Every finding carries a what this does not establish section, and a test asserts the prose never contains malicious, attacker, compromised, breach or exfiltrated. HIVE reports risky system states. It does not accuse anyone.
What I learned
A demo you cannot replay is not evidence. Chasing byte-identical output forced out three separate sources of hidden non-determinism I would otherwise have shipped, each of which would have made a reproducibility claim false.
Detection was the easy half. The hard part was proving containment did not break the thing it was protecting — which turned out to need a counterfactual, a shared apply path, and a verification pass that distrusts the planner.
Negative tests carry more weight than positive ones. A rule that fires on the happy path proves very little. A rule that goes silent when each of its four preconditions is removed, and fires when all four are present, proves something.
Stating limits is load-bearing. The most useful section of the README is What HIVE does not do. Two rules, not all emergent behaviour. Simulated controls, not enforcement. Paths, not payloads. Being precise about the edges made the claims inside them credible.
What's next for HIVE
More rules, built the same way — each with its manifest-deviation gate and its set of necessary-condition tests. Real enforcement adapters behind the existing control interface, so a simulated block becomes a service-mesh policy or a gateway rule. Ingest from real agent frameworks rather than fixtures. And promoting immunity patterns from a single-estate hypothesis toward something that can be evaluated across estates — carefully, because a pattern is a hypothesis about structure, not a proven universal rule.
Log in or sign up for Devpost to join the conversation.