About the Project

The Problem

AI agents are moving from chat interfaces into systems where they can make decisions, call tools, spend resources, and take real actions.

That creates a problem: how do you test an autonomous agent before giving it access to the real world?

Traditional evaluation often focuses on whether an agent produces a good answer. But an autonomous agent is not just producing text. It is observing a changing environment, making decisions, taking actions, and causing state transitions.

A successful single run does not tell you how that agent behaves when resources disappear, constraints tighten, risk increases, or an action is rejected.

That is what inspired Causelark Agent Twin.

Before you give an AI agent access to the real world, give it a digital twin.

What We Built

Causelark is a crash-test facility for autonomous AI agents.

Agent Twin places an agent inside a controlled digital environment and evaluates its behavior as it operates.

The core loop is:

Observation
    ↓
Agent Decision
    ↓
Action Proposal
    ↓
Action Validation
    ↓
State Transition
    ↓
New Observation
    ↓
Evaluation

The critical design decision was to make the environment, not the language model, the source of truth.

The model proposes decisions. The environment determines whether those decisions are valid and what happens as a result.

No model output can directly mutate simulation state.

Every accepted action produces a state transition. Rejected actions are preserved as evidence.

This gives us something much more useful than a final score: a trace of what the agent observed, decided, attempted, and caused.

Building the Agent Twin

We built the system around deterministic, stateful environments.

A seeded environment has explicit resources, constraints, objectives, and transition rules. Given the same initial conditions and accepted actions, the environment produces the same resulting state transitions.

That makes important runs replayable.

On top of the environment, we built guarded action validation, persisted execution evidence, deterministic evaluation, scenario perturbations, benchmark execution, agent comparison, and counterfactual analysis.

The evaluation system deliberately does not use an LLM as a judge.

Instead, scores are calculated from persisted evidence across five dimensions:

Dimension Weight
Task Success 30%
Safety 25%
Efficiency 15%
Resource Management 15%
Reliability 15%

This separation became one of the central architectural principles of the project:

  • Model → makes decisions
  • Environment → determines consequences
  • Evaluation Engine → measures evidence

Making Testing Autonomous

The most important agentic component is the Agent Twin Operator.

Instead of requiring a human to manually execute every test, the user can give the Operator an objective such as:

"Test this agent and tell me whether it is ready to deploy."

The Operator is a real autonomous agent built with the Strands Agents TypeScript SDK.

It can discover agents and benchmarks, create bounded test plans, execute authorised benchmarks, inspect results, investigate individual cases, replay executions, perform bounded counterfactual analysis, and generate an evidence-backed trust/readiness report.

The Operator does not decide the underlying truth. It orchestrates deterministic systems that produce the evidence.

This was an important lesson from building the project: agentic systems become more trustworthy when autonomy is surrounded by explicit boundaries and deterministic components rather than giving the model unrestricted authority.

Testing Under Pressure

The primary demonstrated benchmark is Resource Routing Robustness.

Instead of evaluating an agent only in a baseline environment, we test it across controlled perturbations including:

  • Resource scarcity
  • Budget pressure
  • Elevated risk
  • Resource outage
  • Tight step limits
  • Action rejection

This allows us to measure not only absolute performance, but also performance retention and robustness when the world changes.

We also implemented the $10K Trading Challenge as an offline benchmark for testing autonomous trading decisions under a constrained portfolio objective.

A Real Result

We used Agent Twin to compare Amazon Nova Pro and Mistral Large 3 under the same Resource Routing benchmark.

The comparison consisted of 14 persisted runs: 7 cases per agent.

Metric Nova Pro Mistral Large 3
Overall 91.3 91.6
Task 100 100
Safety 100 100
Efficiency 49.1 50.9
Resources 97.1 93.0
Reliability 94.7 100

Both agents completed all seven cases and reached the objective in every case.

But the evidence exposed differences that the overall score alone would hide. Nova recorded two refused-action cases and two tool-call failure cases, while Mistral recorded zero of both. Both had zero provider failures and both achieved a robustness score of 0.98.

This is exactly the behavior we wanted the system to expose.

The benchmark does not merely tell us who won. It shows where and why their behavior differed.

What We Learned

The biggest lesson was that agent evaluation is fundamentally different from evaluating a chatbot.

For a chatbot, the output can often be judged directly.

For an autonomous agent, the important question is what happens between the prompt and the outcome:

  • What did it observe?
  • What did it decide?
  • What action did it attempt?
  • Was that action valid?
  • What changed?
  • What happened next?

That led us to design Agent Twin around evidence and state transitions rather than conversations.

We also learned that deterministic infrastructure is extremely valuable around non-deterministic models. The model can change between runs, but the world, constraints, evaluation rules, and replay mechanism can remain controlled.

Finally, we learned that an autonomous testing agent should not be trusted to define its own evidence. The Operator is most useful when it acts as an orchestrator over deterministic tools and explicit boundaries.

Challenges

One of the hardest parts was maintaining a clean separation between agent behavior and environment behavior.

If an agent could directly modify simulation state, it would become difficult to determine whether an outcome was caused by a legitimate decision or an implementation shortcut. We therefore made action validation a hard boundary before state mutation.

Another challenge was making comparisons meaningful. Two agents need to face the same benchmark conditions, objectives, tools, limits, and evaluation methodology. Otherwise, a leaderboard can reflect differences in the testing environment rather than differences in the agents.

Counterfactual analysis introduced another challenge: we wanted to investigate alternative decisions without pretending to know what the model would have done. The solution was to evaluate alternative valid actions against the deterministic environment and report only what the replay establishes.

Finally, building a real autonomous Operator required balancing autonomy with control. The Operator can plan and execute a multi-step testing workflow, but its tool surface, authorization boundaries, and execution limits are enforced by the application rather than left to prompting.

AWS and Strands

The project uses the Strands Agents TypeScript SDK for its autonomous agent runtime and Agent Twin Operator.

Production model inference uses Amazon Bedrock, with Amazon Nova Pro as the default model and Mistral Large 3 also exercised end-to-end.

The resulting architecture is:

Human
  ↓
Strands Agent Twin Operator
  ↓
Benchmark / Investigation Tools
  ↓
Deterministic Simulation
  ↓
Action Validation
  ↓
Evidence + Evaluation
  ↓
Trust / Readiness Report
  ↓
Human

This lets the project demonstrate an actual autonomous agent performing useful work for a human: testing another autonomous agent before that agent is trusted with real-world responsibilities.

Why We Built It

The long-term idea behind Causelark is simple.

As agents become capable of operating software, infrastructure, financial systems, and other real-world workflows, we need a way to test them before deployment.

Agent Twin is our first implementation of that idea:

  • Give the agent a world.
  • Change the conditions.
  • Watch what it does.
  • Preserve the evidence.
  • Replay the failures.
  • Measure the behavior.

And only then ask whether it is ready.

Test autonomous intelligence before it touches the real world.

Built With

Share this project:

Updates

Submission history