The Inspiration In an economy of AI agents deciding where to spend credits, the statement "trust me, I reviewed it" is completely worthless. Nobody can verify a review's quality inside a fast-paced market round. A counterexample, however, is entirely different: it either reproduces or it doesn't.

We needed a way to cryptographically prove that an agent's claim about a piece of code (e.g., "this patch fixes the race condition") is accurate, without requiring anyone to trust the LLM's prose.

What it does An agent mid-task submits an artifact and the claim it's making about it. Crucible then actively tries to break the claim:

  1. It generates real attack hypotheses.
  2. It actually executes them in an isolated sandbox.
  3. It only ever reports falsified once it has watched a reproduction fail for real.

The receipt that comes back proves—from the cryptographically signed audit trail SharedOS's kernel wrote while we worked—that we never touched anything outside the artifact we were granted.

How we built it Crucible is built on top of SharedOS and is structured across three deliberately separate planes:

  • Kernel plane: Powered by @aicoo/sharedos, every read of the artifact goes through SharedOSKernel.invokeTool. Access is decided against a purpose-bound CapabilityGrant that is strictly restricted to the job's path and a 5-minute TTL.
  • Adversary plane: A bounded agent driver reads the artifact under the kernel, then shells out to a headless Claude Code CLI for the actual falsification reasoning. Because Claude runs outside the kernel mediation without tools, it cannot touch governed resources.
  • Proving plane: We use a fresh temporary directory and node --permission to create a strict sandbox (filesystem confined, child_process and worker threads denied). A hypothesis is never reported as falsified until the script actually fails in this environment.

At the end of the job, Crucible generates an ed25519-signed receipt containing a hash-chain over the job's exact audit slice.

Challenges we ran into

  • Secure Sandboxing: Node's permission model handles filesystem and process spawning well, but network egress remains a challenge we had to carefully design around.
  • Cryptographic Audit Trails: Wiring the kernel against the actual @aicoo/sharedos SDK (not a mock) to ensure all decisions were written to an audit trail and folded into a verifiable hash-chain without impacting the adversary's latency.

Accomplishments that we're proud of

  • It's real and working end-to-end today. We aren't mocking the system; it genuinely falsifies real TOCTOU race conditions and ownership-check bugs.
  • The node --permission sandbox genuinely blocks out-of-scope access.
  • We successfully built an independently-verifiable receipt system. Anyone can fetch the public key, recompute the hash chain, and know exactly what Crucible did (how many tool calls, if any were writes, or if any reached outside the grant) without having to trust us.

What we learned Isolating the "reasoning" (Claude CLI) from the "resource plane" (SharedOS Kernel) is incredibly powerful. By pulling the LLM entirely out of the authorization loop during the generation of the hypotheses, we drastically reduced the attack surface while still allowing the LLM to write brilliant adversarial tests.

What's next for Crucible

  • Multi-language provers: Currently, the reproduction script is JavaScript. We want to support Python, Go, and Rust artifacts by running their respective toolchains inside equivalently locked-down sandboxes.
  • Network Namespace Isolation: Mitigating the sandbox network egress gap with an OS-level network namespace or an allowlist proxy before handling untrusted artifacts at scale.
  • SharedNet Registration: Finalizing our SharedOS Cloud registration so any agent on SharedNet can dynamically find and call Crucible during market rounds.
  • Escalation Resolution: Wiring up the sharedos.escalate loop so that pending escalations can seamlessly resolve into follow-up grants.

Built With

Share this project:

Updates

Submission history