Inspiration

Every data platform team has lived this incident: a dashboard breaks, someone gets paged, and the fix that ships is "re-run the job closest to the alert." Three hops upstream, the table that actually caused the problem keeps quietly rotting, because nobody walked the lineage graph far enough to find it. DataHub already has that graph. The upstream edges are sitting right there. But today, walking it during an incident is a manual, tab-switching exercise: open the lineage tab, click through each hop, eyeball freshness, guess.

A short deterministic loop, a framework-enforced approval gate, and a graded eval run against an answer key the agent couldn't see. That discipline, a small number of steps done rigorously, every claim backed by evidence a human can independently re-check, tested against both a broken case and a healthy case, and a pass rate reported honestly instead of asserted. That is what we built Root-Cause Tracer around from day one.

What it does

Root-Cause Tracer takes a broken dataset's URN and walks DataHub's real upstream lineage graph, hop by hop, until it finds the deepest node that still shows an anomaly: a freshness gap, a null-rate spike, a schema drift, or a failed job. It doesn't stop at the first broken-looking table it sees; it keeps walking upstream past every node that's still anomalous, because the visibly broken table downstream is usually just the last domino, not the first.

Every claim it makes is backed by a citation: the exact URN, field, and value pulled from DataHub that produced the verdict, plus an executable reproduction query a human can run themselves to independently confirm the anomaly before trusting the diagnosis at all. For the null-rate scenario, it goes one step further and drafts an actual fix diff from the dataset's real transformation query, and it's honest about when it can't: if the root cause is an orchestration failure rather than a code bug, there's no query to fix, so it correctly returns no diff and hands the human a diagnosis report instead of inventing a change.

None of this reaches production DataHub, and no GitHub PR opens, until a human clicks Approve. That gate isn't a system prompt asking the model to be polite about it. It's a structural ApprovalGate in code (src/human_gate.py) that blocks the write-back and PR-generation code paths entirely until an explicit /approve call resolves it. Reject, and the decision is logged with nothing written anywhere.

How we built it, and why each piece exists

The deterministic core: no LLM anywhere near the diagnosis

src/traversal.py walks the graph. src/root_cause_locator.py decides which node is the root cause. Neither file makes a model call. The root cause is whichever upstream node is the deepest one still showing an anomaly signal, computed with plain comparisons against real values pulled from DataHub. This was a deliberate constraint, not a limitation: an LLM asked to "find the root cause" from a lineage graph will happily produce a plausible-sounding wrong answer, and a wrong answer that reads confidently is more dangerous than no answer. So the two model-touching files in the entire codebase, src/narrator.py and src/fix_drafter.py, are given the deterministic conclusion after it's already been reached and are explicitly instructed not to re-diagnose, only to narrate the reasoning in plain English or draft code against a stated root cause. If you strip Gemini out of Root-Cause Tracer entirely (no GEMINI_API_KEY), the diagnosis is unchanged. You just lose the narration and the fix draft.

DataHub is the spine, not a data source you skim

src/mcp_client.py is the single choke point for every read and write against DataHub. Nothing else in the codebase talks to the graph directly. It calls DataHub's real lineage, freshness, null-rate, and job-status metadata through mcp-server-datahub over SSE, using the official MCP Python SDK, with an automatic fallback to the DataHubGraph SDK if the MCP server is unreachable. That fallback matters more than it sounds: a diagnosis tool that goes dark the moment its MCP server hiccups isn't one a platform team would trust on-call.

Critically, Root-Cause Tracer doesn't just read the graph and throw the answer away in a chat window. On approval, it writes back: a tag and a diagnosis document land on the real root-cause entity in DataHub (src/writeback.py), so the next engineer who opens that dataset in DataHub sees the diagnosis sitting on the node itself, not buried in a Slack thread. draft_fix() additionally depends on a DataHub capability the fallback SDK can't provide, get_dataset_queries, so the auto-fix path is only available when the real MCP server is live, which was itself a forcing function to get the MCP integration genuinely working end-to-end rather than treated as optional.

Trust guarantees you can check, not take our word for

We built the eval against planted ground truth the agent has no way to see. scripts/seed_scenarios.py emits a small, purpose-built lineage graph directly onto a live local DataHub instance: real entities, real lineage edges, and a genuinely buggy transformation Query entity, covering three cases:

Scenario Anomaly What the agent must do
Freshness bug Upstream job failure Walk past the broken-looking downstream table to the real staging failure
Null-rate bug Bad transformation query Locate the root cause and draft a real fix diff from the real query
No-action trap Nothing actually wrong Refuse to diagnose a dataset tagged to look suspicious

src/verifier.py grades every run against the planted answer key, and src/eval_harness.py reports a pass rate, false-positive rate, and trap-refusal rate, rendered live in the app's own Eval Results panel (GET /eval), not just written in a README. Current result, graded against real DataHub metadata on a live instance: 3/3 correct, 0/3 false positives, 1/1 traps refused.

The no-action trap is the piece we'd point a skeptical judge to first. It's easy to build an agent that finds a root cause; it's much harder to build one that says "there is no root cause here" when the dataset merely looks suspicious, and that refusal is graded, not assumed.

The approval gate, structurally

src/human_gate.py implements ApprovalGate as an actual state machine (PENDING -> APPROVED / REJECTED), sitting between diagnosis and any side effect. write_diagnosis_to_datahub() and open_pr() are unreachable from PENDING. The block is in the control flow, not in the instructions to the model, so there's no prompt-injection or hallucination path that skips it.

Frontend: a real workspace, not a chat box

React 18 + Vite, no router. A landing page and a five-panel workspace (Diagnose, Lineage Graph, Evidence Log, Approvals, Eval Results) so a judge (or an on-call engineer) can see the lineage graph with the root cause highlighted, the cited evidence and reproduction SQL, the fix diff when one exists, and the full history of past approve/reject decisions, all without leaving the tool. The Approvals panel in particular makes the human gate visible, not just enforced. You can see every diagnosis that was rejected and nothing happened, which is the point.

Open Source: Claude Code Skill

The deterministic diagnosis algorithm is also packaged as a standalone Claude Code Skill — a reusable set of instructions anyone can drop into their own project to get the same evidence-cited, human-gated root-cause tracing directly in a Claude Code session, without running this app at all. It just needs mcp-server-datahub connected as an MCP server against their own DataHub instance.

To use it in another project:

mkdir -p .claude/skills/datahub-root-cause
curl -o .claude/skills/datahub-root-cause/SKILL.md \
  https://raw.githubusercontent.com/Akshat74747/Root-Cause-Tracer/main/skills/datahub-root-cause/SKILL.md

Then just ask Claude something like "why is my_final_table showing unexpected nulls?" and it will follow the same upstream-traversal, evidence-citation, and approval-gate discipline as the app itself.

Challenges we ran into

The hackathon's suggested seed datapacks don't exist. The brief referenced nyc-taxi and healthcare DataHub datapacks; datahub datapack list against a real instance shows only bootstrap and showcase-ecommerce are actually published. Rather than fake it, we wrote scripts/seed_scenarios.py to emit real entities and real lineage edges via the DataHub Python SDK emitter directly, which turned out better for us anyway, since it gave us exact control over the planted anomalies our answer keys depend on.

Deciding what the LLM was allowed to touch. The instinct in an agent hackathon is to let the model do more. We went the other way: traversal and root-cause selection stayed 100% deterministic, and we drew a hard line that Gemini only narrates or drafts against an already-reached conclusion. That constraint is why the eval harness can grade the diagnosis deterministically at all. A nondeterministic root-cause step would make "3/3 correct" a much weaker claim.

Keeping the MCP integration load-bearing instead of decorative. It would have been easy to read from DataHub once via the SDK and call it done. Making mcp-server-datahub the primary path, with TOOLS_IS_MUTATION_ENABLED=true required for write-back, the SSE endpoint quirks, and a real fallback for when it's down, meant the MCP server actually has to be running and actually has to work for the full loop (including get_dataset_queries for fix drafts) to succeed.

Grading honestly instead of asserting. It's tempting to just say "it works." Building verifier.py and eval_harness.py against planted, hidden ground truth, including a trap the agent has to actively refuse, meant the pass rate is something a judge can re-run (python scripts/run_eval.py) rather than something we're asking them to believe.

Accomplishments we're proud of

  • A root-cause locator that's provably deterministic: same graph state in, same root cause out, every time, with the reasoning auditable in plain comparisons rather than a model's hidden weights.
  • Graded 3/3 correct, 0/3 false positives, 1/1 trap refused against real DataHub metadata on a live instance, reproducible with one command.
  • A structural approval gate, not a prompted one. Writing to DataHub and opening a PR are both unreachable code paths until a human explicitly approves.
  • Real write-back to the graph, not just reads: the diagnosis and a tag land on the actual root-cause entity in DataHub, so the tool contributes evidence back to the context graph instead of only consuming it.
  • A fix-drafting path that knows its own limits. It only drafts a diff when a real buggy query exists to draft one from, and honestly returns "no fix, here's the report" when the root cause is an orchestration failure instead.

What's next for Root-Cause Tracer

  • Multi-anomaly correlation: right now each traversal locates one deepest anomalous node; a real incident often has two independent upstream faults, and we'd like the locator to report both rather than stopping at the first.
  • Broader anomaly types: schema drift and column-level lineage are partially modeled; extending to SLA breach and access-pattern anomalies would widen real-world coverage.
  • Contributing scenario packs back to DataHub: the seeded lineage graphs we built for the eval harness (freshness bug, null-rate bug, no-action trap) are generic enough to be useful to anyone building or testing lineage-aware tooling against DataHub, and we'd like to upstream them as a reusable demo datapack.
  • Slack/PagerDuty approval surfaces, so the human gate lives in the tools an on-call engineer is already watching, rather than requiring a trip to the dashboard.

Built With

Share this project:

Updates

Submission history