Inspiration
Different targets can still share one environment.
Two agent actions can each be locally valid, touch different resources, and still compose into an invalid joint outcome because they interact through a constraint neither action owns. That is the gap Interlock addresses.
A per-target lock is useful for the hazard it can see: same-target contention. But a composition hazard can span two distinct targets while remaining coupled through shared environment state. Interlock adds revision-bound environment evidence to the coordination decision before mutation.
What Interlock does
Interlock is a deterministic composition-safety layer, not another agent that recommends caution in prose.
It:
- receives independently valid application-agent intents;
- reads revision-bound WorkspaceJSON environment evidence before mutation;
- evaluates the intents against their shared environmental constraints;
- emits a deterministic coordination decision such as
ALLOW_PARALLELorWITHHOLD_SERIALIZE; - produces an authorization receipt on the protected path;
- requires that receipt at the target so the coordination step is load-bearing rather than decorative;
- separates a mutation's
EXECUTEDreport from an independently authenticatedOBSERVEDread-back; - publishes frozen judge views and commit-pinned evidence so the proof can be inspected without trusting the UI.
The goal is not to make agents less autonomous. It is to make autonomous actions composable and legible before they change shared state.
Measured operational-utility proof
HAC-343 is a controlled local evaluation over a frozen corpus. It compares four mechanically distinct coordination strategies.
| Strategy | Invalid coupled outcomes | Safe parallel opportunities retained |
|---|---|---|
| Uncoordinated | 2/2 | 2/2 |
| Global serialization | 0/2 | 0/2 |
| Per-target locking | 2/2 | 2/2 |
| Interlock | 0/2 | 2/2 |
The per-target baseline is deliberately credible rather than weakened for comparison. It correctly serialized 2/2 same-target contention cases and parallelized 4/4 cross-target cases. Its limitation appeared only when the demonstrated shared constraint crossed two distinct target keys: it missed 2/2 of those bounded composition hazards.
That is the specific distinction Interlock is designed to address. The claim is not that "locks do not work." The lock worked for the hazard it could see. The demonstrated composition hazard crossed its visibility boundary.
Evidence ablation
The evaluation also tests whether the environment evidence is actually load-bearing.
With coupling evidence present, Interlock produced 0/2 invalid outcomes in the coupled conditions. In the perturbed fixtures, the intents and final tree remain fixed while the available coupling evidence changes. When that evidence is removed, the decision reverses to ALLOW_PARALLEL and 2/2 invalid outcomes return.
That is the causal signature: changing the available environment evidence changes the coordination decision and the resulting outcome.
This evaluation ran locally. It did not run on Google Cloud.
Separate Google Cloud participation proof
The cloud proof is a separate recorded traversal with separate evidence.
The recorded path is:
gemini-3.5-flash → Google ADK 1.35.1 / Vertex AI → Cloud Run-hosted agent → Interlock MCP proxy → Interlock decision + authorization receipt → protected mutation EXECUTED → independently authenticated read-back OBSERVED alpha=45 → Cloud Logging correlation.
EXECUTED and OBSERVED remain separate facts. One is what the mutation path reported. The other is what a separately authenticated principal read afterward.
Three negative controls were recorded on the cloud path:
- forged identity header →
403; - invalid bearer token →
401; - direct target bypass without receipt →
403.
Those are three recorded controls, not comprehensive security coverage.
The cloud run proves real participation by Gemini 3.5 Flash, Google ADK, Vertex AI, Cloud Run, the Interlock protected-action path, independent read-back, and Cloud Logging correlation. It does not claim that the local four-arm evaluation was reproduced in Google Cloud.
How I built it
The Interlock composition engine is a deterministic TypeScript core. It accepts revision-bound environment evidence plus a set of agent intents and returns a coordination decision with attribution.
The core carries no presentation dependency. Judge-facing surfaces consume frozen generated evidence projections so proof values are derived from the evidence package rather than manually typed into screenshots.
The demonstrated stack includes:
- Gemini 3.5 Flash for the recorded agent run;
- Google ADK 1.35.1 as the agent framework;
- Vertex AI as the model access path;
- Cloud Run for the recorded Google Cloud execution path;
- Cloud Run IAM / service-account transport for the demonstrated transport-authentication boundary;
- Cloud Logging for run-correlated runtime evidence;
- TypeScript / Node.js for the deterministic Interlock core and supporting services;
- Model Context Protocol (MCP) for the Interlock proxy boundary;
- WorkspaceJSON as revision-bound descriptive environment evidence;
- Vercel for the public judge-facing narrative and verification experience.
The architecture deliberately keeps provenance layers distinct. Google Cloud transport authentication establishes the caller on the demonstrated path; Interlock establishes the application-level coordination decision and authorization receipt. WorkspaceJSON describes the environment but does not authorize anything.
Agent Runtime, Agent Gateway, Memory Bank, Agent Identity, Model Armor, and CONTENT_AUTHZ are not presented as participants in the recorded deployment path.
Judge experience
The public experience is designed in three layers:
- Understand — see why different targets can still share one operational constraint.
- Believe — compare uncoordinated execution, global serialization, credible per-target locking, and Interlock; then inspect the evidence-ablation result.
- Verify — enter the cockpit to inspect frozen evidence, methodology, immutable references, limitations, and raw proof.
The Google Cloud traversal is shown only after a hard proof-class reset so the controlled local evaluation and the deployment proof cannot silently collapse into one stronger claim.
Verify it yourself
Live project: https://interlock.marcellelabs.io/
Repository: https://github.com/Marcelle-Labs/interlock
Core reproduction:
pnpm install
pnpm run check
pnpm run typecheck
pnpm run build
pnpm test
The repository also contains the frozen experiment packets, deterministic aggregators, public cloud evidence, judge exports, architecture assets, cockpit, claim boundaries, and package gates used to keep factual surfaces tied to their source evidence.
Challenges I ran into
Proving evidence changes the decision
A UI can make evidence look important without proving that it affected anything. The ablation arm exists specifically to falsify that. Remove the coupling evidence while holding the task structure fixed and the decision changes; the invalid joint outcome returns.
Comparing against a credible alternative
A global lock is easy to beat on concurrency and therefore not enough. The evaluation includes a real per-target-lock baseline that correctly handles same-target contention. Its bounded failure is narrower and more informative: cross-target composition hazards can sit outside a per-key discipline's visibility.
Keeping proof classes separate
The controlled local evaluation and Google Cloud traversal prove different things. The project deliberately resets context between them rather than letting proximity imply that the four-arm counterfactual ran in the cloud.
Making the receipt load-bearing
Generating a receipt is not sufficient. The protected target must care. The recorded direct-bypass control returns 403 when the demonstrated mutation path omits the required receipt.
Separating self-report from observation
A mutation returning success is not the same as independently observing resulting state. The cloud path uses a separately authenticated read-back so EXECUTED and OBSERVED remain distinct claims.
Accomplishments
- Built a new deterministic composition-safety layer during the hackathon period.
- Compared Interlock against uncoordinated execution, global serialization, and credible per-target locking on a frozen bounded corpus.
- Recorded 0/2 invalid coupled outcomes with 2/2 safe parallel opportunities retained for Interlock in the primary bounded comparison.
- Demonstrated that removing coupling evidence reverses the decision and restores 2/2 invalid outcomes.
- Ran a real Gemini 3.5 Flash + Google ADK + Vertex AI + Cloud Run traversal through the Interlock protected-action path.
- Made the authorization receipt load-bearing at the protected target.
- Captured three cloud negative controls:
403,401,403. - Separated mutation execution from independently authenticated read-back.
- Published commit-pinned evidence and a judge-facing verification surface rather than asking judges to trust screenshots alone.
What I learned
The difficult part of multi-agent coordination is not detecting that two agents exist. It is establishing when independently reasonable actions interact through shared state, making the coordination decision attributable to current evidence, and keeping the proof surface honest about what each run established.
I also learned that observability and provenance cannot be bolted on after the demo. If a judge cannot tell which system asserted a fact, which evidence changed a decision, and whether an outcome was reported or independently observed, the architecture is not yet legible enough.
Limits
The claims are intentionally bounded:
- the HAC-343 controlled evaluation did not run on Google Cloud;
- the recorded cloud traversal does not reproduce the four-arm local counterfactual;
ALLOW_PARALLELis a coordination decision, not a general verification or safety guarantee;OBSERVEDmeans independently reread, not universally safe;WITHHOLD_SERIALIZEis not human approval or a generalized authorization lifecycle state;- Agent Runtime, Agent Gateway, Memory Bank, Agent Identity, and Model Armor did not participate in the recorded path;
- no exactly-once, restart-safety, recovery, universal collision-prevention, fleet-scale, comprehensive security, or production-readiness guarantee is claimed.
What's next
The next step is to broaden the evaluated composition-hazard families and operational constraints while preserving the same evidence discipline: credible baselines, frozen metrics before results, reproducible packets, explicit negative findings, and a hard separation between causal evaluation and deployment proof.
Built With
- gemini-3.5-flash
- google-adk
- google-cloud-iam
- google-cloud-logging
- google-cloud-run
- model-context-protocol-(mcp)
- node.js
- typescript
- vercel
- vertex-ai
- workspacejson