UNWIND — Governed Autonomous Operations

Inspiration

Giving an AI agent more autonomy also gives it more ways to be wrong — in ways nobody notices until the mistake becomes expensive.

A chatbot that hallucinates can give you a bad paragraph. An agent that hallucinates and then acts can revoke access, file a correction, or spend a budget — turning an incorrect decision into a real incident.

That made us ask a different question:

Not just “Can an agent do the task?” — but “What happens around the task?”

Who delegated the work? What evidence did the agent rely on? What happens when two sources disagree? Did a human actually see the consequential action? Can the system recover safely? And does the next mission remember what went wrong?

UNWIND was built around that problem.

Its execution loop is:

objective → plan → delegate → execute → verify → recover → reconcile → govern → record knowledge → evaluate → evolve → next mission

We believe this is a governance problem before it is an intelligence problem.

What We Built

UNWIND runs missions, not conversations.

A mission starts with an operational objective and moves through a stateful, multi-stage execution pipeline. The objective determines the plan and specialist delegation rather than following one fixed script.

Each specialist agent has its own identity, scope, and permissions. Agent outputs pass through deterministic output contracts that validate schema, safety, policy, self-consistency, and evidence grounding before they can influence mission state.

When evidence conflicts, UNWIND does not silently choose one source. A separately-scoped Reconciler evaluates the contradiction using an authority ladder. If the independent derivation disagrees with the first interpretation, that disagreement becomes an explicit state instead of being hidden.

Before consequential actions, the system applies bounded authority, warrant economics, verification, and a mandatory human governance boundary. Mutating routes remain authenticated, while the public Judge Demo uses a separate read-only deterministic replay so judges can inspect the system without weakening the real security boundary.

The Core Innovation: Governance as an Execution Loop

The interesting part of UNWIND is not simply adding security around an agent.

Governance is part of the mission itself.

A worker can fail. A result can violate its contract. Evidence can disagree. Risk can increase. A human can reject an action. Any of these can change what happens next.

For example:

worker result → contract check → reject → REPLAN

or:

conflicting evidence → reconciliation → uncertainty increases → human gate

The system therefore does not treat governance as a final approval screen. It continuously influences execution.

Memory That Changes the Next Mission

UNWIND also has a bounded cross-mission memory loop.

Completed missions distill what they actually measured — including settled or disputed premises, faults, isolations, and evidence coverage — into provenanced knowledge records.

The next mission retrieves only the relevant subset within a bounded context budget.

This is deliberately not unrestricted self-learning.

Recalled knowledge can increase the risk floor or require additional verification, but it cannot grant wider authority, introduce a new tool, or bypass a governance gate.

In our verified scenario, recalled evidence changed the next mission's risk profile from:

LOW → MEDIUM

and that scrutiny was applied directly to the new plan.

This makes memory operational: the past can change how the next mission behaves without becoming a hidden source of authority.

Failure and Recovery

Failure is a first-class state in UNWIND.

Tool execution uses bounded timeouts and retries with explicit failure types such as:

TIMED_OUT · RAISED · CONTRACT

A worker can therefore return something that looks structurally valid but contains fabricated or ungrounded information — and the contract layer can reject it before it reaches mission state.

A faulted step can trigger REPLAN, while unresolved faults cannot silently become a successful mission result.

This was one of the most important design decisions: recovery should be part of autonomous execution, not something operators have to reconstruct afterwards.

Evidence, Provenance and the Mission Time Machine

Every mission stage is persisted as a checkpoint.

The Mission Time Machine reconstructs the actual mission history from those checkpoints, allowing the operator or judge to inspect the mission arc, individual checkpoint state, decisions, and provenance.

This means the system is not merely showing a live animation of an agent.

It can answer:

What happened? Where did it happen? Which agent acted? What evidence existed at that point? What decision followed?

The history is therefore part of the system's operational state, not just a UI log.

Evaluation and Evolution

UNWIND evaluates trajectories, not just outcomes.

Completed missions are scored against seven deterministic behavioral criteria that measure whether the agent remained governed while completing its objective.

The evolution loop can compare agent-version behavior and prevent promotion when a version gains throughput by sacrificing safety or governance.

Importantly, this is not autonomous model retraining. The system governs which agent-version instruction is promoted rather than changing model weights.

What We Learned

The biggest lesson was that governance cannot be bolted onto an autonomous system at the end.

Typed output contracts had to become hard gates rather than optional validation.

Stateful execution had to become durable and monotonic so that replanning could not silently corrupt the mission queue.

Evidence needed provenance — a knowledge record without its mission and checkpoint context is just a string.

Conflicting information needed an independently-scoped second derivation rather than a simple tie-breaker.

And memory needed structural limits so that learning from previous missions could increase scrutiny without ever becoming a new source of authority.

We also learned that honesty about system boundaries is itself an engineering feature. The public Judge Demo is explicitly a deterministic replay of a captured mission trace, while privileged mutations remain authenticated. Features that are architectural but not exercised live are labeled that way instead of being presented as something they are not. :contentReference[oaicite:1]{index=1} :contentReference[oaicite:2]{index=2}

Challenges

One of the hardest parts was making multiple independently-scoped agents communicate through typed contracts without making the contracts either too strict or too permissive.

We also had to make the mission state durable across replanning. An early implementation could renumber steps after a replan and silently point queue entries at the wrong step, which had to be fixed.

Reconciliation was another challenge. We did not want to claim that the mechanism generalized based on one example, so we tested it against a second unrelated evidence scenario with a different dispute path.

Cross-mission recall required another constraint: previous knowledge must be able to increase scrutiny, but it must never be able to manufacture authority. That limitation was therefore designed structurally rather than left as a convention.

Finally, exposing the system to judges without weakening authentication required two separate doors: a credential-free, read-only Judge Demo and an authenticated operator path.

What Makes UNWIND Different

Most agent demos optimize for:

“Can the AI do the task?”

UNWIND optimizes for:

“Can the agent pursue the task, prove what happened, recover safely when something fails, reconcile conflicting evidence, preserve what it learned, and remain accountable to a human?”

That is the system we wanted to build.

Not unrestricted autonomy.

Governed autonomy — with evidence, recovery, memory, verification, and accountability built into the execution loop.

Built With

  • fastapi
  • gemini
  • gemma
  • google-adk
  • google-artifact-registry
  • google-cloud
  • google-cloud-build
  • google-cloud-firestore
  • google-cloud-pubsub
  • google-cloud-run
  • google-cloud-secret-manager
  • google-genai
  • html5
  • javascript
  • lyria
  • opentelemetry
  • opentelemetry-sdk
  • playwright
  • pydantic
  • pytest
  • pytest-asyncio
  • python
  • uvicorn
  • veo
  • vertex-ai
Share this project:

Updates

Submission history