-
-
01-moirae-safe-work-overview-synthetic.png
-
03-exact-human-approval-synthetic.png
-
02-model-proposes-protocol-decides.png
-
08-live-bedrock-structured-call.png
-
09-live-bedrock-call-accounting.png
-
Shows the happy-path product result while clearly preserving the synthetic boundary.
-
10-real-provider-attempt-evidence.png
-
12-mp09-adversarial-acceptance.png
-
11-unknown-blocked-no-retry.png
-
13-mp09-attacks-refused.png
Moirae Protocol Inspiration
We started Moirae with a simple but pressing question:
What happens when an AI agent is allowed to do real work, but we do not want the model itself deciding what it is allowed to do?
As modern agents become increasingly capable of planning, using tools, calling APIs, and completing long-running tasks, a dangerous new problem emerges. The exact same model that proposes an action should not automatically become the authority that approves it, executes it, or declares that it succeeded.
Our earlier project, Moirae Console, explored this through a governed WebMCP workflow. Red-team testing exposed several vulnerabilities that became the foundation for this new protocol: approvals could be too transient, validation could become optional, provider calls could bypass policy boundaries, and a successful API response could easily be mistaken for proof of a real-world effect.
Moirae Protocol is our answer. It is a system where the model can propose actions, but deterministic software and independent governance actually decide what happens next.
The core rule is simple:
The model proposes; the deterministic Protocol and Fates decide.
What it does
At its core, Moirae takes an agent's proposal and passes it through a strictly governed execution pipeline.
A model might propose something like sending appointment details, rescheduling a meeting, or transmitting a customer contact directory. However, we never treat the model's output as the final authority.
Instead, the proposal goes through this journey:
Validated against a strict schema. Compiled into a deterministic ActionIntent. Submitted to Fates for independent admission. Approved automatically, blocked, or routed to a human. Persisted through a locally durable approval boundary. Placed onto a locally durable work queue. Claimed by a bounded worker. Executed through the governed provider boundary. Observed independently. Reconciled by Horae. Projected into a simple human-facing state.
To the end-user, this complex machinery is distilled into four understandable outcomes:
Handled automatically Needs you Blocked Activity
Crucially, these categories are not guessed by the browser or the model.
For example, Handled automatically only appears when durable execution is complete and Horae has enough trusted evidence to support a confirmed result.
We built this on the philosophy that approval alone is not enough. A queue item is not enough. An API success response is not enough. Even an SES MessageId is not enough.
We demand evidence.
How we built it
We built Moirae Protocol primarily in TypeScript and Node.js, using the Strands Agents SDK with Amazon Bedrock for the agent boundary.
Rather than generating arbitrary executable instructions, the agent produces a strictly structured proposal, which is then fed into our deterministic pipeline.
Identity is carried through the entire workflow rather than reconstructed later. The original request, deterministic ActionIntent, approval, queue work, claim, execution attempt, provider correlation, and reconciliation evidence remain bound through stable identifiers and hashes.
This prevents an approval for one action from being replayed or accidentally applied to another.
Agent layer
The Strands integration uses @strands-agents/sdk with Amazon Bedrock to produce structured AgentProposalV1 output under bounded turns, token limits, explicit schema validation, and no automatic inference retry strategy.
We exercised this boundary live against Amazon Bedrock with three bounded inference calls.
All three returned structured output that passed schema and semantic validation, with:
three calls authorized; three calls attempted; zero retries; zero extra probes.
Crucially, a successful model response still has no authority.
The proposal remains an untrusted input until deterministic Protocol code compiles it into an exact ActionIntent and Fates independently decides what may happen next.
Governance layer
The heart of our governance is Fates, backed by our Ananke work.
Fates independently evaluates a deterministic ActionIntent to decide whether it is allowed, denied, or requires human approval.
The model cannot override this decision.
For actions needing human judgement, we built a locally durable approval system that tightly binds the approval, the ActionIntent, the exact target, and the trusted host context.
Approval is therefore not vague permission to "continue." It is authority for one exact action identity.
Background execution
Approved work enters a locally durable queue with deterministic identities, lease and claim semantics, bounded retries, and restart-oriented recovery.
Even then, the worker receiving a queue item does not gain the authority to declare that the external effect actually happened.
Execution authority and effect truth remain deliberately separate.
Effect reconciliation
This is where Horae steps in.
Moirae distinguishes between:
making a provider request; receiving a provider response; and proving that the intended external effect actually occurred.
Fates and Horae are deliberately separated from the model-facing agent layer.
Fates governs whether work may proceed.
Horae governs what can truthfully be said about the result.
For our AWS SES path, we built a governed SES v2 adapter using:
a verified dedicated sender; a verified sandbox recipient; a fixed template; trusted server-side configuration; a dedicated least-privilege send role; deterministic correlation tags; an SES configuration set; Amazon SNS; and an Amazon SQS observation queue.
This allows provider-event evidence to be observed independently from the original SES request.
A provider response is therefore not treated as proof that the real-world effect occurred.
We exercised that boundary once against the real SES provider path.
Moirae made exactly one governed provider invocation, with automatic retries disabled.
No independently correlated SES event arrived and no provider operation ID was available, so Horae classified the outcome as:
UNKNOWN
Moirae did not report success.
It did not retry the action.
The attempt was permanently consumed, the work remained reconciliation-required, and the product remained blocked.
That result became one of the most important proofs in the project: Moirae can distinguish between attempting an action and knowing what actually happened.
Product layer
Finally, our MP-07 product projection translates durable protocol truth into the four judge-facing categories:
Handled automatically Needs you Blocked Activity
We intentionally kept the frontend free of policy engines and protocol classifiers.
It simply renders the trusted product view.
If an auditor wants to dig deeper, technical evidence such as approval IDs, work IDs, execution IDs, and reconciliation state remains available through progressive disclosure.
The primary UI therefore stays understandable without hiding the receipts.
The judge-facing local interface is intentionally synthetic so that the full approval and happy-path UX can be demonstrated safely without producing an external effect.
The real Strands/Bedrock inference evidence and the governed provider experiment are preserved separately as live evidence.
Live, synthetic, and confirmed evidence are deliberately kept distinct so the demo UI cannot accidentally overstate what happened.
Challenges we faced
One of our biggest realizations was that the most dangerous mistakes usually happen at the boundaries between components, not inside the model call itself.
A system can have excellent model prompts and still be fundamentally unsafe if:
approval is treated as execution; queue delivery is treated as authority; provider acceptance is treated as effect truth; or an unknown result is silently converted into success.
We had to make these distinctions explicit architectural invariants.
Durability was another major hurdle.
Human approvals, work queues, execution state, and reconciliation evidence must survive normal process boundaries without allowing authority to drift between actions.
Otherwise, a system may behave correctly during a short demo while losing the evidence needed to govern long-running autonomous work.
On the practical side, AWS integration threw several curveballs at us.
We worked through account authentication, SES sandbox requirements, verified identities, IAM scoping, and independent delivery observation.
During our first Provider-02B preflight, we discovered that the AWS SDK's default SES retry behavior could permit multiple network attempts.
Rather than ignore it, we stopped the live run.
We reconfigured the governed SES client to explicitly use:
maxAttempts: 1
and added offline behavioral tests proving that even retryable failures result in exactly one underlying request-handler invocation.
That became a defining lesson for the project:
Safety guarantees should be demonstrated, not assumed from defaults.
Adversarial validation
We also treated the protocol itself as something that needed to be attacked.
MP-09 exercised 15 focused adversarial classes, including:
model-authority injection; browser-authority injection; approval replay; cross-attempt authority reuse; recipient substitution; duplicate execution; stale claims; false provider confirmation; mismatched observation evidence; and illegal promotion of UNKNOWN into success.
The focused suite passed 247 tests with no accepted P0, P1, or P2 findings.
We then gave the frozen implementation to Claude and DeepSeek for separate adversarial reviews.
Those reviews did identify hardening issues, including a real weakness where synthetic observation evidence could satisfy the live reconciliation path.
We fixed the submission-relevant issues and ran targeted retests.
Both targeted external retests passed with no reproducible critical bypass.
We do not present this as a formal security certification. It is adversarial engineering evidence: attack, identify weaknesses, remediate, and retest.
What we learned
The biggest takeaway from building Moirae is that reliable agent systems require much more than clever prompt engineering.
They require explicit boundaries between:
proposal authority approval execution observation truth
We also learned to embrace UNKNOWN as a useful and legitimate result.
If an external action might have happened but cannot be proven, the safest response is not to blindly retry or claim success.
It is to preserve the uncertainty and reconcile it.
We realized that human-in-the-loop design is much more nuanced than simply adding an Approve button to a UI.
Human approval should grant eligibility for one exact action.
It should not bypass policy, perform the action itself, or allow the frontend to manufacture a successful result.
Finally, our earlier red-team work on Moirae Console proved invaluable.
The attacks and vulnerabilities discovered there became explicit foundational rules in the Protocol, including:
MODEL_OUTPUT_NOT_AUTHORITY
APPROVAL_NOT_EXECUTION
BROWSER_STATE_NOT_AUTHORITY
PROVIDER_SUCCESS_NOT_EFFECT_CONFIRMED
UNKNOWN_NOT_CONFIRMED
SYNTHETIC_NOT_LIVE
Where it is now
Moirae Protocol is now a fully integrated submission build with the core governed-agent path implemented and publicly reproducible.
We have completed:
Live Strands/Bedrock inference: three bounded real inference calls produced structured, schema-valid proposals with zero retries. Deterministic proposal compilation: model output is converted into an exact ActionIntent before authority is considered. Independent Fates admission: actions are allowed, blocked, or routed to explicit human approval outside the model. Exact-action human approval: approvals are bound to the action, target, trusted host context, and presentation identity rather than acting as general permission. Locally durable work and execution state: queue, claim, approval, attempt, and reconciliation identities are preserved through the governed execution workflow. Governed Provider-02B execution: the real SES provider boundary was invoked exactly once with automatic retries disabled. Independent effect reconciliation: because no independently correlated SES event was observed, Horae classified the result as UNKNOWN; MP-06 remained RECONCILIATION_REQUIRED, MP-07 remained BLOCKED, and resend was prohibited. Judge-facing product projection: the local interface presents trusted protocol state as Handled automatically, Needs you, Blocked, or Activity while clearly separating synthetic demonstration data from live evidence. Adversarial acceptance: fifteen focused attack classes were exercised, followed by external Claude and DeepSeek adversarial reviews and targeted retesting. Public reproducibility: the frozen repository is public and its GitHub Actions CI is green against the accepted Fates and Horae materializations.
What we do not claim is a production-hosted runtime or confirmed delivery of the live SES effect.
Hosted durable infrastructure and production deployment remain future work.
That distinction is intentional.
Moirae is designed not merely to make agents act, but to preserve the difference between:
being allowed to act attempting the action and knowing what actually happened
Ultimately, that represents the broader vision behind Moirae:
An autonomous system should not merely perform actions — it should be able to explain exactly who authorized them, what happened, and why the system believes that result is true.
Built With
- agentic-ai
- ai-agents
- amazon-bedrock
- amazon-ses
- amazon-sns
- amazon-sqs
- amazon-web-services
- human-in-the-loop
- node.js
- responsible-ai
- strands-agents-sdk
- typescript
Log in or sign up for Devpost to join the conversation.