Inspiration

I'm a trained epidemiologist with about 10 years in pharma supporting drug development. When we approach the development of a study we begin with a directed acyclic graph (DAG).

In a causal DAG, every node and arrow expresses a scientific assumption. Those assumptions determine which biases may be present, which variables should or should not be adjusted for, and whether a causal effect is identifiable.

We created DAG Studio to give epidemiologists and other researchers an approachable, browser-based environment for constructing and examining those assumptions. However, an AI assistant interacting with a conventional web application still has to interpret what it 'sees' on the canvas. It cannot reliably distinguish a parts of a graph from a decorative line or know whether its internal description of the graph still matches what the researcher sees.

WebMCP suggested a better model: give the agent structured access to the same scientific object the researcher is editing.

That inspired DAG Studio WebMCP, a shared causal-modeling workspace for researchers.

What it does

DAG Studio WebMCP allows a researcher and an AI agent to collaborate on the same live causal DAG. The researcher sees and edits the visual canvas, while the agent uses nine structured WebMCP tools connected to the exact same client-side state:

  • get_current_dag
  • add_node
  • remove_node
  • add_edge
  • remove_edge
  • analyze_current_dag
  • check_adjustment_set
  • generate_analysis_code
  • simulate_current_dag

An agent can read the current graph, propose changes, inspect open backdoor paths, identify minimal sufficient adjustment sets, verify a researcher-proposed adjustment set, and generate R or Python analysis code.

The simulation tool uses DAG Studio's existing linear Gaussian structural equation model to generate synthetic data from the live graph. For each node (X_j), the illustrative model follows

[ X_j = \sum_{i \in \mathrm{pa}(j)} \beta_{ij}X_i + \epsilon_j. ]

The tool returns bounded summaries, correlations, a small data preview, model-implied total effects, and crude and minimally adjusted ordinary least squares estimates. In our demo, a researcher can ask the agent to retain results from two different DAG hypotheses in the conversation and compare their implications using the same sample size, random seed, and coefficients.

These simulations do not determine which DAG is scientifically correct. They reveal what follows from the assumptions encoded in each model.

How we built it

DAG Studio WebMCP extends the existing React and Vite DAG Studio application. We deliberately reused the established causal and simulation functions in dag-engine.js and its dag-engine.d.ts declarations rather than rewriting the algorithms for the Challenge.

The WebMCP layer registers tools through the current document.modelContext API. It also includes a compatibility fallback for hosts implementing the earlier navigator.modelContext location.

The WebMCP tools receive dependency-injected access to the same synchronous state references and React mutation functions used by the visual canvas. Therefore:

  • an agent mutation immediately appears on the canvas;
  • a human canvas mutation is immediately visible to the next tool call;
  • agent and human graph edits share the same undo history; and
  • analysis and simulation always operate on the graph currently visible to the researcher.

Agent changes receive visible attribution. The simulation tool also opens DAG Studio's existing simulation interface, so the result is reflected in both the conversation and the application.

Everything remains browser-side. The Challenge version is an isolated static HTTPS deployment on Cloudflare Pages, separate from the production DAG Studio website. The source is published under the MIT license, and GitHub Actions runs the original engine-parity suite together with the new WebMCP shared-state tests.

Challenges we ran into

The most important challenge was not tool registration, rather it was guaranteeing that the agent and researcher were truly modifying the same graph. React state updates are asynchronous, while a sequence of agent tool calls may need to observe each mutation immediately. We addressed this by wiring the tools to the canvas's synchronous references and existing mutation wrappers rather than creating a second “agent graph.”

WebMCP is also an emerging standard, and browser implementations differed during development. The API appeared under different objects across environments, some Chrome testing calls required JSON-serialized arguments, and not every ChatGPT or Codex browser session exposed site tools. Clear connection diagnostics and compatibility support were essential.

We also had to preserve the scientific boundary between assistance and authority. A variable being associated with an outcome does not automatically make it a confounder, and simulated data cannot validate a biological hypothesis. Tool descriptions, responses, and visible language therefore characterize graph changes as assumptions or proposals for researcher review.

Finally, simulation results can become too large for a useful agent interaction. We bounded sample-size options, limited row previews, summarized variables and correlations, validated coefficient overrides, and made every run reproducible with a recorded random seed.

Accomplishments that we're proud of

  • Nine WebMCP tools operate on the exact graph visible in DAG Studio.
  • Human and agent mutations remain synchronized and undoable.
  • The original DAG Studio causal algorithms were reused rather than approximated by the language model.
  • The agent can analyze competing hypotheses and retain reproducible simulation results for comparison.
  • Simulation outputs explicitly separate model-implied effects from crude and adjusted estimates.
  • The integration includes current and earlier WebMCP API compatibility and visible connection diagnostics.
  • Thirteen automated tests cover engine parity, tool registration, shared state, mutations, cycle rejection, adjustment analysis, code generation, simulation reproducibility, input validation, bounded results, visible simulation state, and a two-hypothesis comparison.
  • The project is available as an open-source repository and a static public HTTPS application.

What we learned

WebMCP is most valuable when it exposes meaningful domain operations rather than reproducing every button in an interface. analyze_current_dag is more reliable and scientifically legible than asking an agent to visually inspect a set of arrows. simulate_current_dag is more useful than teaching an agent how to navigate a simulation panel one control at a time.

We also learned that shared state changes the nature of human–AI interaction. The canvas is no longer just something the agent can see; it becomes a common scientific object that both participants can inspect and revise. The conversation can retain multiple hypotheses and their results even though the visual editor displays the current model.

Most importantly, greater agent capability makes explicit epistemic boundaries more important. The engine can verify the graphical and statistical consequences of a DAG, but neither the engine nor the agent can decide whether its assumptions are scientifically true.

What's next for DAG Studio: Agent-Ready Causal Modeling with WebMCP

Next, we want to deepen hypothesis-centered workflows: named DAG versions, persistent side-by-side comparisons, researcher approval checkpoints, and exportable analysis reports that preserve the connection between assumptions, adjustment decisions, simulations, and generated code.

We also see opportunities for collaborative provenance, where every proposed node or edge can carry a rationale, citation, reviewer comment, and acceptance status. Additional simulation families could support binary, count, survival, longitudinal, and nonlinear outcomes while preserving transparent assumptions about the data-generating process.

The long-term goal is not an autonomous system that declares causal truth. It is an agent-ready scientific workspace that helps researchers articulate alternatives, expose disagreements, test implications, and document why a final model was chosen. Ideally this workspace can be called upon in a multi-agent orchestration that runs a research workflow from end to end.

Live application: https://dag-studio-webmcp-sandbox.pages.dev/

Source code: https://github.com/Black-Swan-Causal-Labs/dag-studio-webmcp

Original DAG Studio: https://dagstudio.blackswancausallabs.com/

Built With

Share this project:

Updates

Submission history