Inspiration
I build tools for domains where "I don't know" has to be a first-class answer. My previous work is in variant classification: deciding whether a specific mutation found in a patient's genome actually causes disease. The field's standard guidelines define five verdicts, and one of them is Variant of Uncertain Significance, an explicit reportable class meaning the evidence does not support a call in either direction. It is not a failure to classify. It is a classification, and a clinician must never confuse we could not determine this with this is fine. That distinction ended up shaping every design decision in this project, including the one I am most confident about: this tool will not write "safe" into a graph.
The concrete starting point was DataHub's own guidance. The platform documents that column-level lineage is what target leakage detection requires, and that a review of the lineage graph is how a leak gets found before a model reaches production. The verb is review. DataHub provides the resolution; a person does the looking. That is the gap.
There is a second question underneath it, and it is the one I could not put down. A feature can be clean for the model it serves today and unsafe for a different one. Feature stores do not record what target a feature was trained against, so nobody checks this. Answering it needs the label inside the graph, next to the features, which is exactly what was missing.
What it does
Two components, one story.
An emitter writes ML lineage from training code into DataHub: the feature list, each feature's source column, the label column, the training run, the deployment. These are facts that exist in a practitioner's training script and are recorded nowhere. y = df["churned"] is a declaration of the target that no system captures.
A scanner and agent walks the filled graph and looks for leakage at column resolution. Two detectors: direct derivation, where a feature descends from the label column itself, and shared-ancestor contamination, where a feature and the label descend from a common column. A separate staleness check reads the Timeline API for upstream schema change since training.
Findings are written back into the graph, not into a report: a tag on the model, a structured property carrying the detector list and score, the full evidence path as a note, and an incident raised on the root-cause dataset with its owner resolved. A later reader, human or agent, inherits the reasoning, and can ask the catalog which entities this agent has written to.
And it answers the second question as a matrix: every feature against every label, not just the model each feature currently serves. On the demo estate that surfaced a feature that is clean for the model it serves and would be contaminated if reused for a different one.
How I built it
Three layers, with a hard boundary between them.
The graph layer is fully deterministic: traversal, ancestor-closure intersection, budgeted expansion, evidence-chain construction. No LLM touches it. The semantic layer is the only place an LLM appears, and it never invents lineage. It chooses among candidates the deterministic layer produced. The write layer is deterministic again.
Reads go through DataHub's MCP Server, and the detectors are exposed as tools designed to compose with the Agent Context Kit's own, demonstrated in examples/agent_tools_demo.py. Writes are split: tags and structured properties through MCP, entity creation through the SDK, incidents through GraphQL, because that is where each capability actually lives.
The part that took longest was not the detector. It was measurement. Eight leakage classes taken from published work, with a visibility prediction for each committed to git before the generator that tests them. Randomized estates. A three-arm ablation that drops column lineage per model. An adversarial pass where an LLM was given the detector's description and asked to defeat it. And a run against a real third-party dbt project, because an estate I built myself cannot tell me what I do not already know.
Accomplishments that I'm proud of
Not the detector. The detector is the easy part.
A correction of mine is now in DataHub. The MLflow connector's documentation described lineage coverage it does not produce. I measured it against a registered model, filed the correction, had one of my claims corrected by a maintainer in review, verified what remained, and it was merged into master during this hackathon. Five more contributions are open.
Two strangers independently reproduced bugs I reported. One confirmed the cached-empty-lineage family on a different DataHub version through a different client. Another found the same telemetry stall from a different side, before the server answers initialize at all. I did not ask either of them. That is the difference between a report and a finding.
The tool still refuses to certify. This was the constraint I set at the start, and it was under pressure the whole way, because a governance tool that never says "safe" is harder to demo. It held. safe_for_labels does not exist in the output, and the reason it does not exist is measured rather than asserted.
What I learned
The finding I did not expect, and the one I would keep if I kept nothing else:
Certification is anti-monotonic in graph completeness. An incomplete metadata graph does not merely limit what a tool can verify; it inflates what the tool believes it has verified.
Every implementation error that shrinks the ancestor closure produces a false certificate. None produces a false "unknown." Errors decay toward confidence.
I found this three times. A metric that counted a finding as a certification. An ingestion arm holding less metadata that certified more, because the columns that would have produced gaps were absent from its schema entirely, so the walk terminated cleanly on a smaller world. And a traversal bug that stopped one hop early on a mirror edge.
Then I measured it. A correction that could have moved verdicts either way moved 36 to "incomplete," 1 to a finding, and 0 to certification.
The design rule follows directly, and it is why safe_for_labels does not exist in this tool's output:
Certification must rest on every edge being resolved, not on no gap being noticed.
Challenges
The metamodel gap. mlFeature.sources accepts dataset-level references only. Column-level provenance is expressible for datasets and not for features, and features are the entity the failure mode runs through. I verified this twice, through the error my own emitter received and by observing DataHub's Feast connector reach the same limit. The tool carries column-level provenance in customProperties as a workaround, and an RFC is open. This is a gap between a vision and its implementation, not a defect, and surfacing it is ordinary contributor work.
Silent wrong answers, in the platform I was building on. The tool exists because data failures do not announce themselves. Building it, I met several instances of the same shape: a lineage query issued while the index was still building, whose empty result was cached and then served, including to the UI. A removed column edge served for a full cache TTL. A documented search idiom that returns rc=0 with zero results when one extra quoting layer is added, indistinguishable from an empty catalog. A delete reporting "Hard deleted 1 entities" for a URN that never existed. Twenty-seven such observations are logged; thirteen went upstream, one is merged, and two were independently confirmed by other builders.
A demo that had to survive unattended. The hosted instance crashed every twelve hours while reporting healthy the whole time. The cause was not memory, which is what I diagnosed first and got wrong: the OpenSearch container's PID 1 is the JVM, its healthcheck runs curl every five seconds, and nothing reaps the children. Seven thousand zombie processes later, thread creation fails. The fix is five characters, init: true, and it is filed.
My own measurements lying to me. Fifteen times, a number that looked clean turned out to be a failure that had not written anything, a control that could not fail, or a counter read right after the event that resets it. Fourteen of them are catalogued. The rule that came out of it is short: a verification is not a verification unless you have written down how it would fail.
What this does not do
Five of the eight leakage classes are defined over rows, values, or training code. None of those exist in a lineage graph, and this tool cannot see them. That prediction was registered before it was tested, and the injected cases in those classes produced zero detections across 54 injections, which is what the prediction said.
They are not gaps to close later. A tool claiming to catch them would be lying about its own abstraction.
What's next for ML Guard
The next steps are mostly other people's decisions, which is the honest position for a tool built on someone else's metamodel.
If the RFC lands, a workaround disappears. Column-level provenance for features currently lives in customProperties because mlFeature.sources cannot express it. If that gap closes upstream, the workaround comes out and the provenance becomes queryable by anything, not just by this tool.
Two known defects are documented and unfixed, deliberately. An incident is deduplicated by title, so it can outlive the finding that raised it. Retirement cannot prove it owns a generic risk tag, so it never removes one. Both are written up with their mechanisms, and both were left alone because fixing them changes what a rescan writes, which would have invalidated every measurement already frozen. They are the first work after submission, not before it.
The CLI needs model-scoped writes. scan takes --model; the agent loop does not. That makes arming writes an all-or-nothing decision over a whole catalog, which is the wrong shape for a governance tool.
What is not next: the five invisible classes. Row-level, value-level and training-code leakage are not in a lineage graph and will not be reached by adding detectors here. Catching them needs a different abstraction and probably a different tool. Saying so is not modesty, it is the same rule the rest of this project runs on.


Log in or sign up for Devpost to join the conversation.