DataPhylax

Three agents that find, fix, and verify data quality problems in a DataHub catalog — coordinating through the catalog itself, on a locally fine-tuned 7B model.


Inspiration

Data catalogs answer "what data do we have?" They don't answer "what's broken?"

That gap is where the expensive failures live. A pipeline stops running and the dashboard keeps showing last week's numbers. Someone renames a column and three models downstream start training on nulls. An engineer leaves and nobody notices they owned forty tables until one of them breaks at 2am.

We found a real example of this in DataHub's own sample catalog while testing. A table called order_details has a description that explicitly says it contains PII and PCI data classifications. Not one of its columns is tagged. Nobody planted that — it's just what happens when documentation and governance drift apart.

The hackathon brief said it plainly: without reliable knowledge of schemas, lineage, ownership and governance, agents hallucinate or get stuck on tasks any data engineer could finish in minutes. We wanted to build the thing that supplies that knowledge, and then acts on it.


What it does

Scout runs six checks across the catalog:

Check What it finds
freshness_violation a table past its refresh SLA
ownership_gap an asset with nobody responsible for it
pii_regression a column that looks like personal data and isn't tagged
silent_duplicate the same table under two names
lineage_break an upstream that no longer exists
schema_drift a column dropped, renamed, or narrowed

The detection is deterministic — regex, set comparison, timestamp arithmetic. No model involved. A regex either matches or it doesn't, and we did not want a language model guessing at whether a column exists.

Every finding then goes to a fine-tuned Qwen2.5-7B running locally, which writes the incident report and judges the severity. Scout writes the result back into DataHub as structured properties.

Fixer reads those incidents back out of DataHub — not from Scout directly. Reversible metadata changes it applies itself: tagging a PII column, flagging an asset for access review, marking a duplicate deprecated. Anything touching production code — migrations, pipeline repairs, ownership — it writes as a proposal file for a human to review.

Historian then checks whether any of it worked. Fixer reporting success only means an API call returned 200, so Historian re-runs the detector that raised each finding and sees whether it's actually gone. It aggregates (event type → fix type → success rate) and writes that back to DataHub, so the next Fixer run knows which remediations are proven and which have never been tried.

The three agents never talk to each other. DataHub is the coordination layer.


How we built it

The model

We needed a model that reliably emits a fixed JSON schema and makes sensible severity calls. Prompting a general model gave inconsistent structure, so we fine-tuned.

Training data. We wrote a scenario generator that synthesises findings across five industry domains (e-commerce, fintech, healthcare, logistics, adtech), four severity levels, varying lineage depth, and four asset types — about 1,400 distinct combinations. Then we distilled 499 (finding → report) pairs from DeepSeek V4 Pro.

Severity was deliberately not an input. We gave the teacher a rubric and let it judge from the evidence, so the student would learn the judgement rather than parrot a label. The generator was weighted so genuinely dangerous situations appeared often enough for critical to be learnable.

Fine-tune. QLoRA on Qwen2.5-7B-Instruct — 4-bit NF4, rank 16, alpha 32, lr 2e-4, three epochs, loss masked so only the report contributes. Best checkpoint was epoch 2 (eval loss 0.5003; epoch 3 rose to 0.5137). About 28 minutes on a single A10G.

Serving. Converted to GGUF, quantised Q4_K_M, served with llama.cpp. That took throughput from 13 tok/s to 76 tok/s — roughly a 6x improvement — and shrank the model from 15GB to 4.4GB.

Evaluation

40 held-out scenarios generated with a fresh seed:

Metric Result
Valid JSON 40/40 (100%)
Schema compliant 40/40 (100%)
Hallucinated identifiers 0/40
Severity — exact match vs teacher 74%
Severity — within one level 97%

The severity number needs context, and this is the part we're most careful about. We ran the teacher model twice over the same 40 inputs at temperature 0. It agreed with itself 82% of the time. Severity is a judgement call, not a fact — so 74% is about 90% of the achievable ceiling, not 74% of a perfect score. We would rather explain that than quote a number that flatters us.

The zero-hallucination result matters more than it looks. Every identifier in the output is checked against the identifiers present in the input. Zero inventions is what makes it safe for Fixer to generate remediation from these reports — a report that references a column which doesn't exist would produce a fix against nothing.

Making it adaptable

Rules get things wrong, and a system that can't be corrected is a system nobody deploys. Two mechanisms, both stored in DataHub rather than in code:

Exceptions. A data owner can suppress a finding ("this replica is deliberate, another team tracks it") or force one the rules missed ("this free-text column contains customer names"). Scoped to a whole asset, a single column, or a specific duplicate pair. A reason is mandatory — an unexplained suppression is a liability six months later.

Configuration. PII patterns, duplicate thresholds, tier rules and freshness defaults all live in DataHub. Built-in defaults mean it works with zero setup; a company overrides only what they care about. Adding acct_no to the bank-account patterns is one CLI call, not a code change.


Challenges we ran into

DataHub's aspect model is stricter than it looks. We spent well over an hour discovering that corpGroup doesn't support institutionalMemory at all, and that corpGroupInfo silently drops customProperties — the write returns 200 and the data simply isn't there. We eventually stored config on a dataset entity instead. This is undocumented, and it's what we want to contribute back.

Writes are asynchronous. A successful ingestProposal means the proposal was accepted, not applied. We chased phantom failures for a while before realising we were reading back before Kafka had processed the write. Several "bugs" were just impatience.

One incident per asset was wrong. Our first write-back stored the highest-severity finding per asset. But customer_features has four problems at once — stale, unowned, untagged PII, broken upstream. Storing one lost three, and Fixer could only ever fix one of them. Moving to an array meant patching every component that read it: Fixer, Historian, and the API each assumed a single object.

Fixer was inventing owners. Our first version guessed a team name from the table name and assigned it. That's worse than leaving the gap visible — now there's a wrong name attached and nobody notices the problem. We changed it to propose an owner inferred from upstream lineage, with the evidence written out, and let a human confirm.

Compute. We started on AMD MI300X and ran out of credits mid-project, so we moved to an AWS g5.xlarge. The 16GB of system RAM then became the constraint — merging the LoRA into the base model in fp16 OOM-killed the box twice. We ended up skipping the merge entirely and letting llama.cpp apply the adapter at load time.


What we learned

Deterministic where correctness matters, model where judgement does. This was the design decision everything else followed from. Detection has to be exact — a regex match is a fact. Severity is a judgement that needs weighing blast radius against asset tier against exposure. Putting the model on the wrong side of that line would have made the system unreliable in the way that matters most.

Verification is the honest part. It would have been easy to have Fixer report its own success. Having Historian independently re-run the detector is what turns "we applied a fix" into "the problem is gone." It also means our knowledge base has rows that say not yet measured rather than 0% — because Fixer only proposed those, and nothing was verified. Showing them as 0% would misrepresent the system.

A teacher's self-consistency is the real ceiling. We nearly retrained on a 74% severity match before thinking to check what the teacher scored against itself. Knowing the ceiling changed the conclusion completely.


What's next

  • Contributing back to DataHub — documentation of the entity/aspect compatibility gaps we hit, which cost us hours and aren't written down anywhere.
  • A second adapter for code generation. Fixer currently templates its proposals. Training on (incident + code file → patch) pairs would let it generate real migrations.
  • Dismissal learning. Exceptions are set manually today. Historian already tracks outcomes; extending it to track rejected findings would let Scout suppress patterns humans keep dismissing.
  • Real-catalog validation. Everything is tested against DataHub's sample catalog plus a fintech ML pipeline we seeded. Running it against a production catalog is the honest next test.

Built With

Share this project:

Updates

Submission history