Inspiration

Production ML incidents rarely start in the model. They start upstream: a column type changes after features were computed, a PII field quietly joins the feature set, a source table goes stale, an owner leaves. Each fact lives in DataHub — but nobody connects them until something breaks.

What it does

Model Lineage Guard is a CLI agent that scans DataHub ML lineage for seven production risks: schema drift after feature computation, PII exposure without an approved exception, stale upstream data, missing ownership, post-outcome feature leakage, model performance regression, and deployment config drift.

Every scan produces a machine-readable JSON report, an interactive HTML report with a severity-colored lineage graph (works fully offline), and a PR-ready Markdown comment. --fail-on high turns any scan into a CI gate that blocks pull requests on lineage risk. With explicit confirmation, the agent writes findings back to DataHub as audit tags through MetadataChangeProposals — dry-run by default, policy-controlled, and logged to an append-only audit trail. Findings that need human judgment are excluded from automated write-back by policy.

The Problem

Production ML incidents rarely start in the model—they start upstream. A column type changes after features were computed. A PII field quietly joins the feature set. A source table goes stale. An owner leaves the team. Each fact lives in DataHub, but nobody connects them until something breaks in production.

What We Built

Model Lineage Guard is a DataHub-aware CLI agent that scans ML lineage for seven production risks that hide between datasets, feature tables, models, and deployments:

  • Schema drift after feature computation
  • PII exposure without approved exceptions
  • Stale upstream data missing expected updates
  • Missing ownership registration
  • Post-outcome feature leakage
  • Model performance regression
  • Deployment configuration drift

Every scan generates three reports:

  • JSON: machine-readable findings and lineage
  • HTML: interactive severity-colored lineage graph (fully offline-capable)
  • Markdown: CI/CD-ready pull request comments

The Safety-First Design

Write-back is explicit and safe by default:

  • --write-back dry-run builds MetadataChangeProposal payloads without sending them
  • --write-back apply --confirm requires explicit human confirmation
  • Findings needing human review are excluded from automation by policy
  • Every decision is logged to an append-only audit trail
  • All changes are tagged back to DataHub for traceability

With --fail-on high, any scan becomes a CI gate that blocks pull requests on lineage risk—shifting ML governance left into the deployment pipeline.

How We Built It

Tech stack: Python 3.11+, Typer, acryl-datahub, datahub-agent-context, Jinja2, vis-network

The architecture is layered:

  1. DataHub access: Lineage walks with graceful fallbacks across server versions, batched aspect fetches
  2. Risk checks: Seven structured checks that turn metadata into findings
  3. Reporting: JSON, HTML, and Markdown artifacts with a dependency-free SVG fallback
  4. Write-back: Policy-controlled MetadataChangeProposal builder with audit logging

CI verification goes beyond dry-run: The build boots a real DataHub instance, seeds demo lineage, discovers models through the Agent Context Kit, scans them, applies write-back with --write-back apply --confirm, and then reads the tags back off the live entities. If a tag is missing, the build fails. This proves tags actually land in DataHub—not just that payloads were built.

Challenges We Solved

The DataHub SDK's scroll_lineage endpoint returns 404 on quickstart GMS, with the error wrapped in an OperationalError that made naive fallback handling dead code. We implemented capability probing on first request, caching the server version, and graceful fallback to the relationships API. This ensures scans work across DataHub versions.

What We're Proud Of

  • A green live-DataHub CI integration pipeline that proves write-back actually lands
  • A write-back design that is safe by default (dry-run mode, --confirm gates, policy files, audit logs)
  • A fully offline demo that works without Docker or internet—and produces the same finding count as the live path
  • Seven production-grade risk checks built with testability and deployment reality in mind

How we built it

Python 3.11 + Typer CLI. Metadata access goes through the DataHub SDK (lineage walks with graceful fallbacks across server versions, batched aspect fetches) and the Agent Context Kit tool surface (mcp_tools.search inside a DataHubContext for model discovery). Reports render with Jinja2 and vis-network, with a dependency-free SVG fallback when there is no internet. CI runs ruff, strict mypy, 72 unit tests, and a live integration job that boots DataHub with datahub docker quickstart, seeds demo lineage, discovers models through the Agent Context Kit, scans them, applies write-back with --write-back apply --confirm, and then reads the tags back off the live entities. A dry run only proves a payload was built; the build fails unless the tag is actually present in DataHub.

Challenges we ran into

The SDK exposes scroll_lineage, but quickstart GMS builds 404 on /openapi/v3/lineage/scroll — and the SDK wraps the HTTP error in its own OperationalError, which made naive fallback handling dead code. We now probe once, remember server capability, and fall back to the relationships API (issue draft ready to file upstream).

Accomplishments that we're proud of

A green live-DataHub integration pipeline, a write-back design that is safe by default (dry-run, --confirm, policy file, audit log), and a demo that works with no Docker and no internet.

What we learned

Building production ML governance requires both technical depth and operational safety. We learned that agents touching metadata write-back must default to dry-run and audit logging—not as afterthoughts, but as core design. We learned that DataHub's metadata model is flexible enough to express complex risk signals through tags, but that graceful fallbacks across API versions are essential when working with community servers. We learned that offline-capable reporting (with SVG fallbacks) removes barriers to adoption in airgapped environments. Most importantly, we learned that teams want to gate ML lineage risks in CI/CD pipelines—the --fail-on feature became central to the project because blocking a PR on upstream risk prevents incidents before they reach production.

What's next for Model Lineage Guard

Per-check threshold tuning: move beyond binary severity to configurable thresholds so teams can tune risk tolerance per check and deployment. More governance checks: extend beyond lineage to include model card completeness, SLA adherence, and custom governance rules. Scheduled mode: keep risk tags continuously fresh by running scans on a schedule and updating DataHub asynchronously. PyPI release: package and release on PyPI so teams can install via pip install mlguard. Integration with DataHub's ML Model Registry features for deeper model-centric checks. Support for write-back of remediation suggestions (not just risk tags) so teams can auto-generate DataHub issues or Slack notifications with actionable next steps.

Built With

  • acryl-datahub
  • actions
  • cli
  • data
  • datahub
  • datahub-agent-context
  • detection
  • docker
  • github
  • governance
  • jinja
  • lineage
  • management
  • metadata
  • ml
  • mypy
  • pytest
  • python
  • risk
  • ruff
  • tracking
  • typer
  • vis-network
Share this project:

Updates

Submission history