Inspiration
ML teams often don't find out a model was trained on stale or unowned data until it breaks in production. By then it's an incident, not a warning. DataHub already stores the lineage that would catch this - training data, feature tables, model relationships - but nobody's watching it automatically. I wanted to build the automated reviewer that actually reads it, on a schedule, and speaks up before things break.
What it does
ML Drift Guardian is an agent that:
- Scans DataHub for ML Model entities and their upstream training data
- Runs detection rules for common silent failures: missing model ownership, unowned upstream datasets, stale training data (not refreshed within a configurable threshold), and missing lineage entirely
- Summarizes findings into a plain-English risk report using an LLM (Groq LLaMA)
- Writes findings back into DataHub as notes on the affected entities, so anyone browsing the catalog sees them
- Sends alerts for critical findings
- Runs on a schedule automatically, and exposes a
/scanendpoint to trigger on demand
How we built it
FastAPI service that queries DataHub's GraphQL API for ML models and their metadata, runs a small rule engine against that data, calls Groq's LLaMA API to turn raw findings into a readable summary, writes results back to DataHub, and persists history (local JSON for dev, DynamoDB-ready for AWS deployment). Packaged with Docker for deployment to AWS free tier (EC2 t2/t3.micro).
Challenges we ran into
- DataHub's
upstreamLineageaspect isn't registered formlModelentities directly in the version I ran locally (it's built for datasets/dashboards) - worked around this by storing upstream dataset references directly in the model'scustomProperties, which turned out to be simpler and just as demonstrable. - Docker Desktop resource allocation on Windows needed bumping up before DataHub's GMS service would respond reliably.
- GitHub's push protection caught a Groq API key that accidentally ended up in
.env.exampleinstead of.env- a good reminder to double check example files before committing, and I rotated the key immediately after.
Accomplishments that we're proud of
Went from zero DataHub experience to a fully working local pipeline: DataHub running locally with seeded ML entities, agent detecting real issues (missing owner, stale data), LLM-generated summaries, and write-back into the DataHub catalog - all functioning end to end.
What we learned
Hands-on experience with DataHub's metadata model and GraphQL API, and how to design an agent that treats a metadata catalog as its source of truth rather than polling application logs directly.
What's next for ML Drift Guardian
- Upgrade from GraphQL calls to DataHub's MCP Server for deeper, more native integration
- Add more detection rules (schema drift on upstream datasets, model staleness relative to retraining cadence)
- Deploy persistently on AWS EC2 free tier with DynamoDB-backed history
- Slack/Teams alerting integration
Log in or sign up for Devpost to join the conversation.