Inspiration

ML teams often don't find out a model was trained on stale or unowned data until it breaks in production. By then it's an incident, not a warning. DataHub already stores the lineage that would catch this - training data, feature tables, model relationships - but nobody's watching it automatically. I wanted to build the automated reviewer that actually reads it, on a schedule, and speaks up before things break.

What it does

ML Drift Guardian is an agent that:

  • Scans DataHub for ML Model entities and their upstream training data
  • Runs detection rules for common silent failures: missing model ownership, unowned upstream datasets, stale training data (not refreshed within a configurable threshold), and missing lineage entirely
  • Summarizes findings into a plain-English risk report using an LLM (Groq LLaMA)
  • Writes findings back into DataHub as notes on the affected entities, so anyone browsing the catalog sees them
  • Sends alerts for critical findings
  • Runs on a schedule automatically, and exposes a /scan endpoint to trigger on demand

How we built it

FastAPI service that queries DataHub's GraphQL API for ML models and their metadata, runs a small rule engine against that data, calls Groq's LLaMA API to turn raw findings into a readable summary, writes results back to DataHub, and persists history (local JSON for dev, DynamoDB-ready for AWS deployment). Packaged with Docker for deployment to AWS free tier (EC2 t2/t3.micro).

Challenges we ran into

  • DataHub's upstreamLineage aspect isn't registered for mlModel entities directly in the version I ran locally (it's built for datasets/dashboards) - worked around this by storing upstream dataset references directly in the model's customProperties, which turned out to be simpler and just as demonstrable.
  • Docker Desktop resource allocation on Windows needed bumping up before DataHub's GMS service would respond reliably.
  • GitHub's push protection caught a Groq API key that accidentally ended up in .env.example instead of .env - a good reminder to double check example files before committing, and I rotated the key immediately after.

Accomplishments that we're proud of

Went from zero DataHub experience to a fully working local pipeline: DataHub running locally with seeded ML entities, agent detecting real issues (missing owner, stale data), LLM-generated summaries, and write-back into the DataHub catalog - all functioning end to end.

What we learned

Hands-on experience with DataHub's metadata model and GraphQL API, and how to design an agent that treats a metadata catalog as its source of truth rather than polling application logs directly.

What's next for ML Drift Guardian

  • Upgrade from GraphQL calls to DataHub's MCP Server for deeper, more native integration
  • Add more detection rules (schema drift on upstream datasets, model staleness relative to retraining cadence)
  • Deploy persistently on AWS EC2 free tier with DynamoDB-backed history
  • Slack/Teams alerting integration

Built With

Share this project:

Updates

Submission history