Inspiration Modern clinical AI systems—like Dr. T’s MedGemma 27B clinical diagnostic models and longevity feature stores—rely on complex data pipelines spanning PostgreSQL, S3, dbt, Airflow, and Feast. When upstream data schemas change, PII tags are missed, or data drift occurs, traditional pipelines break silently. This causes downstream AI models to produce degraded or unreliable medical risk assessments.

We were inspired by DataHub’s open-source Metadata Control Plane (MCP) and Agent Context Kit to build autonomous, metadata-aware AI agents. Instead of treating data governance as a passive documentation exercise, we envisioned active agents that read live DataHub metadata, generate compliant code, and protect production clinical ML models in real time.

What it does DataHub Agents is an end-to-end metadata governance and AI agent control plane built directly into the Dr. T platform. It empowers autonomous data squads through four core capabilities:

Autonomous Data Squads (Agents That Do Real Work): Agents query DataHub’s GraphQL/MCP endpoints to audit schemas, detect missing PII/HIPAA tags, evaluate data quality assertions, and write back governance metadata directly to the DataHub catalog.

Metadata-Aware Pipeline CodeGen: Before writing SQL or Python, agents inspect upstream DataHub schema contracts and lineage graph nodes to generate 100% compliant dbt models, Airflow DAGs, and PySpark ingestion scripts—automatically creating verified GitHub Pull Requests (#142).

Production ML Lineage Protection: The agent continuously monitors end-to-end data lineage from raw clinical telemetry (postgres.drt_clinical_v2) to Feast feature views and MedGemma 27B ML models. If upstream schema drift is detected, the agent triggers an automated circuit breaker to pause vulnerable model endpoints and patch the feature pipeline before bad data corrupts patient risk scores.

Dr. T Ecosystem Bio-Market Governance: Unifies clinical FHIR R4 datasets, Cosmos Green Farm interstellar bio-crop telemetry, and Dr. T Institute transcripts under standardized DataHub metadata contracts.

How we built it DataHub Open-Source Stack: Built on DataHub's metadata architecture (GMS Server, REST/GraphQL MCP APIs, and Agent Context Kit patterns).

React 18 & TypeScript: Designed a high-density, dark-mode control plane UI featuring real-time terminal streaming logs, catalog entity explorers, interactive lineage graphs, and code artifact sandboxes.

Tailwind CSS & Lucide Icons: Styled with sophisticated neutral tones, amber/emerald indicators, and crisp typography to deliver an executive-grade MLOps console.

MedGemma & Feature Store Integration: Linked downstream ML model health metrics (medgemma-27b-longevity-diagnostic) to DataHub upstream lineage dependencies.

Challenges we ran into Bidirectional Metadata Syncing: Enabling AI agents to not only read DataHub entities via GraphQL but also safely write back schema tags, assertions, and incident URNs (urn:li:incident:drt-drift-2026) without race conditions.

Simulating Real-Time Lineage Propagation: Modeling multi-tier lineage (PostgreSQL dbt Feast MedGemma) in a responsive browser interface while maintaining instant visual feedback during simulated data drift events.

Accomplishments that we're proud of Zero-Break Pipeline Generation: Generating dbt SQL and Airflow DAG artifacts that inherit HIPAA masking and DataHub assertions on the very first try.

Automated Circuit-Breaking MLOps: Demonstrating a closed-loop incident handler that catches upstream data anomalies and halts downstream clinical model inference before clinical decisions are impacted.

Unified Catalog Explorer: Providing a single interactive view for dataset fields, PII security flags, quality health scores, and lineage dependencies.

What we learned Metadata as Agent Context: Providing LLM agents with structured DataHub schema contracts drastically reduces hallucinated SQL column names and broken pipeline joins.

Lineage is Crucial for AI Safety: In clinical and longevity research, model governance begins at the raw data ingestion layer, not at the model evaluation phase.

What's next for Datahub Agent Live DataHub Cloud/GMS Webhooks: Expanding real-time Kafka event streaming to trigger instant agent remediation tasks upon any DataHub catalog mutation.

Automated Data Contract Enforcement: Enabling agents to automatically generate and validate JSON Schema / Great Expectations contracts directly from DataHub glossary terms.

Expanded MedGemma Telemetry: Adding automated bias and epistemic uncertainty tracking to DataHub MLModel aspects across multi-hospital clinical trials.

Upgrade datahub agents with thêse: Live DataHub Cloud/GMS Webhooks: Expanding real-time Kafka event streaming to trigger instant agent remediation tasks upon any DataHub catalog mutation.

Automated Data Contract Enforcement: Enabling agents to automatically generate and validate JSON Schema / Great Expectations contracts directly from DataHub glossary terms.

Expanded MedGemma Telemetry: Adding automated bias and epistemic uncertainty tracking to DataHub MLModel aspects across multi-hospital clinical trials.

Built With

  • all
Share this project:

Updates

Submission history