Inspiration

ML models in production don't fail loudly. They fail silently.

When a model breaks it almost never throws an exception — it keeps returning predictions with perfect confidence, and they're just wrong. And the cause is almost never in the model code. It's upstream: a training source added a column or changed a type, a feature's source table got deprecated, an upstream table stopped refreshing three weeks ago, the engineer who owned the pipeline left, or the feature table used at serving time quietly started reading a different source than the one the model trained on.

The entire ML observability market — Arize, WhyLabs, Evidently, SageMaker Model Monitor — watches model outputs. That is a lagging indicator by construction. By the time prediction drift crosses a threshold, you have served bad predictions for days or weeks, and you still have to run a manual investigation to find out why.

Then we looked at what DataHub already stores:

dataset --DerivedFrom--> mlFeature --Consumes--> mlModel --DeployedTo--> mlModelDeployment
   ^                                                 ^
   └──── dataProcessInstance (MLFLOW_TRAINING_RUN) ───┘
         inputEdges = what training ACTUALLY consumed

The complete causal chain is already in the graph. Nothing walks it backwards to flag models at risk before they degrade. That gap is the whole project. Lineage monitoring is a leading indicator; output monitoring is a lagging one.


What it does

Recall walks every production model's full upstream metadata closure in DataHub, evaluates seven detectors over it, scores each model 0–100, and then writes the findings back into DataHub so the next engineer or agent inherits the diagnosis instead of repeating the investigation.

  MODEL                SCORE                 TIER       DEPLOYED
  fraud_detector_v3      100   ██████████    CRITICAL          2
  churn_predictor         48   █████░░░░░    AT_RISK           1

  CRITICAL  Feature 'device_risk_score' reads from a soft-deleted dataset
  CRITICAL  Serving reads 'device_signals_v2', which training never consumed
  CRITICAL  Schema of 'retail.raw.transactions' changed since the last scan
  CRITICAL  'retail.core.fct_orders' has not refreshed in 40 days
  WARN      'retail.core.dim_customers' puts 2 production model(s) at risk

  ───── BLAST RADIUS ─────
  retail.core.dim_customers → 2 production model(s) · 3 live deployment(s)

The seven detectors

# Detector What it catches
1 training_source_schema_drift A dataset the model trained on changed shape — column added, dropped, or retyped
2 orphaned_feature An mlFeature whose source dataset is missing, soft-deleted, or deprecated
3 upstream_staleness A source dataset stopped refreshing past its window
4 ownership_vacuum The model or a critical upstream has no owner — nobody to page
5 training_serving_skew The serving feature table resolves to a different dataset set than the training run consumed
6 undocumented_critical_path A high-fan-in dataset feeding a PROD model with no description and no glossary term
7 blast_radius Inverts the graph: how many PROD models one at-risk dataset endangers, ranked by live deployments

Detectors 5 and 7 don't exist in DataHub or in any output-drift tool we could find.

Detector 5 is detectable precisely because of a gap in DataHub's own graph. MLFeatureTableProperties.mlFeatures is a Contains edge with isLineage: false, so the serving path never appears in a lineage query — you cannot see it by traversing lineage. Recall reconstructs the serving path from raw aspects and diffs it against DataProcessInstanceInput.inputEdges, the ground truth of what training actually read. The set difference is structural training/serving skew: the model is scoring on a data source it was never fitted against. No output-drift monitor sees this, because the predictions look completely normal.

Detector 7 inverts the graph and answers the question an on-call engineer actually has: "this one schema change puts three production models at risk — here they are, ranked by live deployments." One fix clears all of them.

Risk scoring

score = Σ severity_weight(findings) × (1 + 0.15 × active_deployments)   capped at 100

CRITICAL ≥ 60    AT_RISK ≥ 30    WATCH > 0    HEALTHY = 0

Every weight, threshold, and severity lives in recall/recall.yaml — retuning the whole engine never requires a code change.

What it writes back into DataHub

The judging criteria reward going beyond reading metadata. Recall writes back five ways, all idempotent:

Write Target Mechanism
Custom assertion + run event dataset upsert_custom_assertion + AssertionRunEvent aspect
recall.riskScore / recall.riskTier mlModel structured properties
recall:critical / recall:at-risk tag mlModel add_tags
Incident dataset (+ affected models named) raiseIncident GraphQL
Pre-mortem document knowledge base save_document / datahub.sdk.Document

Structured properties are the sleeper. Once Recall writes recall.riskScore, it is queryable through DataHub's own search index, and the tags Recall writes appear as facets on the ML Models browse page. DataHub becomes the risk dashboard. Any tool that can reach DataHub's GraphQL API can ask "which models are at risk" without knowing Recall exists. We deliberately did not build a React dashboard — composing DataHub's shipped surfaces is the more interesting answer, and the more useful one.

Assertion URNs are derived deterministically from (detector, subject), so a re-scan updates the existing assertion rather than piling up a new one every run.


How we built it

The architectural thesis: deterministic Python does the detection; the LLM only explains and remediates.

Graph traversal and rule evaluation are pure functions over metadata — fast, unit-testable, reproducible. A language model traversing lineage would be slow, non-reproducible, and would invent edges that don't exist. So detection is ordinary Python, and every detector is a pure function (closure, config, state) -> [Finding]. Time enters only as closure.scanned_at_ms; history only through an injected state store. An AST walk over recall/detectors/ asserts mechanically that no detector reads the clock, touches the filesystem, or opens a socket — the purity rule is enforced by a test, not by discipline.

The agent's job starts after detection, where the task is genuinely linguistic. recall scan --narrate runs a tool-using Claude agent over the DataHub MCP server. Given the skew finding, it made 4 MCP calls and came back having read the deprecation note naming the replacement table, diffed the two schemas (device_risk → risk_score, plus a new model_version column), noticed the column hadn't merely been renamed but rescaled — "vendor-supplied risk score in [0,1]" versus "rescaled vendor risk score" — identified the single owner who has to make the call, and read the open incidents on the replacement table before recommending a cutover to it. Then it wrote remediation grounded in all of it.

The agent is granted read-only MCP tools. Write-back is deterministic and belongs to writeback/, not to something a model decides mid-sentence.

Three narration tiers, tried in order: the agent (Claude Code credentials, no API key), then a plain LLM call with ANTHROPIC_API_KEY, then deterministic templates. A judge with no Claude credentials at all still gets a complete, accurate pre-mortem. The deliverable is never hostage to an optional dependency.

┌──────────────────────────────────────────────────────────────┐
│  DataHub OSS (localhost)                                     │
│  models · features · feature tables · datasets · runs        │
└───────────┬──────────────────────────────────┬───────────────┘
            │ READ (aspects)                   │ WRITE BACK
            ▼                                  ▲
  ┌───────────────────┐             ┌──────────────────────┐
  │  Graph Resolver   │             │   Writeback Layer    │
  │  (deterministic)  │             │  · custom assertions │
  │  upstream closure │             │  · risk properties   │
  │  per model        │             │  · tags              │
  └─────────┬─────────┘             │  · incidents         │
            │                       │  · pre-mortem docs   │
            ▼                       └──────────▲───────────┘
  ┌───────────────────┐  findings              │
  │  Detector Engine  │───────────┐            │
  │  7 pure functions │           ▼            │
  └───────────────────┘  ┌──────────────────┐  │
                         │ Narrator Agent   │──┘
                         │ (Claude via MCP) │
                         └──────────────────┘

Everything runs against DataHub OSS, locally. seed/ builds a realistic ML estate (3 models, 2 feature tables, 8 datasets, training runs, deployments, owners, glossary terms), break_it.py injects five failures, and reset.py restores it so the demo is re-runnable on camera.


Challenges we ran into

Almost every hard problem here came from the same source: the difference between what the metadata model says and what a running server does. Every one of these was found by measuring, not by reading.

1. Never depend on Elasticsearch during a scan. scrollAcrossLineage, entity search, and timeseries aspects — including Operation, which is the natural freshness signal — are all ES-backed and lag writes. A judge running make seed && make scan back-to-back would get an empty closure, an empty model list, or phantom staleness: failures that look exactly like bugs in Recall.

There turned out to be a second, sharper reason. Against the seeded estate, a raw searchAcrossLineage from fraud_detector_v3 returns the full chain — 3 mlFeature at degree 1, datasets at degree 2–3. But DataHubClient.lineage.get_lineage() returns only the datasets: its hardcoded GraphQL carries inline fragments for Dataset and DataJob only, so ML entities are silently dropped. A closure built on that helper would have been missing the entire serving path — exactly what the flagship skew detector depends on.

So the closure is built from graph.get_aspect() reads, which hit GMS and are immediately consistent; model enumeration comes from a manifest the seed emits; and freshness resolves through a fallback chain (Operation → DatasetProperties.lastModified → custom property). Lineage search is still exercised for the pre-mortem's "prove the path" section, but the scan reports it rather than depends on it.

2. The quickstart server trails the CLI, and it matters. acryl-datahub==1.6.0.15 brings up server v1.5.0.6. Two things break:

  • report_assertion_result() sends an AssertionResultSeverity field the older server's GraphQL schema doesn't define → 500. Recall emits the AssertionRunEvent aspect directly instead; raw aspect emission talks to GMS and isn't coupled to the GraphQL schema at all.
  • raiseIncident rejects mlModel URNs on this version. IncidentInfo.pdl on master does list mlModel in the IncidentOn relationship — but that hasn't reached a release. Recall probes the capability once and degrades to raising the incident on the offending dataset with the affected models named in the description.

The lesson we kept relearning: master-branch source is not evidence of shipped behaviour. We were confidently wrong about this exact thing until the live server corrected us.

3. A JSON patch to structuredProperties writes but never indexes. Patching the risk score is the cleaner write, and it works — the value lands in GMS and reads back correctly through get_aspect. It is simply never indexed for search. The whole point of writing the score is that DataHub's search becomes the query surface, so a patch silently trades away the entire feature and the failure is invisible until you try to query it. Recall writes the full aspect and make verify-search proves the filter returns exactly the at-risk models.

4. The flagship detector's real enemy is false positives. Every supervised training run consumes a labels table that serving has no reason to read. Unfiltered, detector 5 fires on a perfectly healthy estate and instantly loses all credibility. Label-shaped inputs are excluded, the two directions carry different severities (serving-only is CRITICAL, training-only is WARN), and a model with no recorded training run reports an INFO gap rather than a phantom CRITICAL.

5. Escalation has to key off dataset role, not column identity. MLFeatureProperties.sources points at a dataset, never a column. "This column backs this feature" is therefore not computable from the graph, and claiming otherwise would mean inventing a lineage edge DataHub does not have. Recall escalates on whether the changed dataset sits on the serving path or the training path, and attaches name-matching columns as supporting evidence only.

6. Structured properties are not free-text searchable. Typing recall.riskScore > 45 into DataHub's search bar returns nothing — that's a structured filter, not a text query. We verified the mechanism through GraphQL and then wrote it up as UI search-bar syntax, which was simply wrong. The UI shot that actually works is the recall:critical tag facet on the ML Models browse page. Verify claims on the surface you're describing, not on an adjacent one.


Accomplishments that we're proud of

  • Detector 5 exists at all. Finding a real failure mode that is invisible to DataHub's own lineage queries because of a documented isLineage: false edge, and then reconstructing it from aspects, is the piece of this we'd defend hardest.
  • 444 tests, no DataHub required. Three carry outsized weight: all 7 detectors × a healthy estate → zero findings (the false-positive net — every misfire caught here is one that would otherwise happen live, on camera); all 7 detectors × an empty closure → no exception (demo survival); and the AST purity walk.
  • The write-back is real, and it's five paths. Assertions, structured properties, tags, incidents, and knowledge-base documents. The metadata graph is strictly richer after a scan than before it.
  • An upstream contribution we actually needed. See below.
  • It degrades gracefully at every seam. No Claude credentials, no MCP server, an older GMS that rejects mlModel incidents, a first run with no schema baseline — each produces a correct, clearly labelled result instead of a stack trace.

What we learned

DataHub's ML entity model is considerably richer than its ML tooling — the mlModel / mlFeature / mlFeatureTable / mlPrimaryKey / mlModelDeployment / dataProcessInstance graph can express training-versus-serving provenance completely, and almost nothing reads it that way. Most of the interesting work in this project was figuring out which edges carry isLineage: true and what that implies about what you can and cannot see through a lineage query.

The other lesson is methodological, and we relearned it three times: verify on the surface you intend to describe. We verified the risk-score filter via GraphQL and documented it as search-bar syntax; verified incident targeting from the PDL on master and documented it as shipped behaviour; verified a patch write via get_aspect and assumed it was indexed. All three were wrong in the same shape. Every claim in the README and in the demo script is now something that was actually run on the surface being described.


Open-source contribution

docs/advanced/patch.md in datahub-project/datahub states that patch-builder support is wanted for "Containers, Data Flows (Pipelines), Tags, Glossary Terms, Domains, and **ML Models." There is no MLModelPatchBuilder today, which means writing MLModelProperties clobbers the whole aspect — the exact hazard Recall hits when it annotates a model.

We needed it, so we built it — and found the gap is deeper than the docs imply. A JSON patch is applied server-side by a registered Java Template, so a Python builder alone produces well-formed MCPs that GMS then rejects with "template" is null. Measured against v1.5.0.6, patching an mlModel:

Aspect Patchable
globalTags, ownership, glossaryTerms, structuredProperties ✅ entity-agnostic templates are registered
mlModelProperties ❌ no template — 500

So the contribution is two halves, both staged in scripts/upstream_pr/: the Python MLModelPatchBuilder (with tests) and the server-side MLModelPropertiesTemplate that makes it work. Verified live: a patch adding a tag to a model left its description, version, 3 features, 3 metrics, and 2 deployments untouched.


What's next for Recall

  • Ship the two-half patch PR upstream and close the gap patch.md names.
  • Column-level closure. fineGrainedLineages on UpstreamLineage would let detector 1 escalate on the column that actually backs the feature rather than on dataset role — turning today's supporting evidence into the primary signal.
  • Continuous mode. The scan is idempotent and fast (3 models, 0.13s, 8 batched aspect calls), so it wants to be a scheduled job that raises an incident the moment a model crosses into CRITICAL, rather than something you run by hand.
  • Recall as an MCP server. The scan results are already structured; exposing them as tools would let any agent ask "which of my models are at risk and why" and get the causal chain back.
  • More detectors from the same graph: feature-table freshness versus model SLA, glossary-term/PII propagation into feature sets, and deployment-versus-registered- version mismatch.

Try it

Requires Python 3.10+, Docker with ~8 GB available, and uv.

make install     # virtualenv + dependencies
make up          # DataHub OSS locally — http://localhost:9002 (datahub / datahub)
make seed        # a healthy ML estate: 3 models, 2 feature tables, 8 datasets
make scan        # baseline — everything HEALTHY
make break       # five realistic failures
make scan        # every one of them caught

make demo runs the whole sequence unattended; make reset restores the estate. make verify proves every write-back path against your instance and prints which capabilities it has. make test runs all 444 tests with no DataHub required.

The five injuries are all things that happen on a normal Tuesday. Two of them are halves of one story: somebody began migrating device_signals to a v2 table, pointed the serving feature at the new one, retired the old one — and never retrained the model. That single half-finished migration leaves the feature reading a dead table and a table the model was never fitted against.

Built With

  • acryl-datahub
  • claude
  • datahub
  • datahub-mcp-server
  • datahub-oss
  • docker
  • graphql
  • pytest
  • python
  • pyyaml
Share this project:

Updates

Submission history