DataPhylax
Three agents that find, fix, and verify data quality problems in a DataHub catalog — coordinating through the catalog itself, on a locally fine-tuned 7B model.
Inspiration
Data catalogs answer "what data do we have?" They don't answer "what's broken?"
That gap is where the expensive failures live. A pipeline stops running and the dashboard keeps showing last week's numbers. Someone renames a column and three models downstream start training on nulls. An engineer leaves and nobody notices they owned forty tables until one of them breaks at 2am.
We found a real example of this in DataHub's own sample catalog while
testing. A table called order_details has a description that explicitly
says it contains PII and PCI data classifications. Not one of its columns is
tagged. Nobody planted that — it's just what happens when documentation and
governance drift apart.
The hackathon brief said it plainly: without reliable knowledge of schemas, lineage, ownership and governance, agents hallucinate or get stuck on tasks any data engineer could finish in minutes. We wanted to build the thing that supplies that knowledge, and then acts on it.
What it does
Scout runs six checks across the catalog:
| Check | What it finds |
|---|---|
freshness_violation |
a table past its refresh SLA |
ownership_gap |
an asset with nobody responsible for it |
pii_regression |
a column that looks like personal data and isn't tagged |
silent_duplicate |
the same table under two names |
lineage_break |
an upstream that no longer exists |
schema_drift |
a column dropped, renamed, or narrowed |
The detection is deterministic — regex, set comparison, timestamp arithmetic. No model involved. A regex either matches or it doesn't, and we did not want a language model guessing at whether a column exists.
Every finding then goes to a fine-tuned Qwen2.5-7B running locally, which writes the incident report and judges the severity. Scout writes the result back into DataHub as structured properties.
Fixer reads those incidents back out of DataHub — not from Scout directly. Reversible metadata changes it applies itself: tagging a PII column, flagging an asset for access review, marking a duplicate deprecated. Anything touching production code — migrations, pipeline repairs, ownership — it writes as a proposal file for a human to review.
Historian then checks whether any of it worked. Fixer reporting success only means an API call returned 200, so Historian re-runs the detector that raised each finding and sees whether it's actually gone. It aggregates (event type → fix type → success rate) and writes that back to DataHub, so the next Fixer run knows which remediations are proven and which have never been tried.
The three agents never talk to each other. DataHub is the coordination layer.
How we built it
The model
We needed a model that reliably emits a fixed JSON schema and makes sensible severity calls. Prompting a general model gave inconsistent structure, so we fine-tuned.
Training data. We wrote a scenario generator that synthesises findings across five industry domains (e-commerce, fintech, healthcare, logistics, adtech), four severity levels, varying lineage depth, and four asset types — about 1,400 distinct combinations. Then we distilled 499 (finding → report) pairs from DeepSeek V4 Pro.
Severity was deliberately not an input. We gave the teacher a rubric and
let it judge from the evidence, so the student would learn the judgement
rather than parrot a label. The generator was weighted so genuinely dangerous
situations appeared often enough for critical to be learnable.
Fine-tune. QLoRA on Qwen2.5-7B-Instruct — 4-bit NF4, rank 16, alpha 32, lr 2e-4, three epochs, loss masked so only the report contributes. Best checkpoint was epoch 2 (eval loss 0.5003; epoch 3 rose to 0.5137). About 28 minutes on a single A10G.
Serving. Converted to GGUF, quantised Q4_K_M, served with llama.cpp. That took throughput from 13 tok/s to 76 tok/s — roughly a 6x improvement — and shrank the model from 15GB to 4.4GB.
Evaluation
40 held-out scenarios generated with a fresh seed:
| Metric | Result |
|---|---|
| Valid JSON | 40/40 (100%) |
| Schema compliant | 40/40 (100%) |
| Hallucinated identifiers | 0/40 |
| Severity — exact match vs teacher | 74% |
| Severity — within one level | 97% |
The severity number needs context, and this is the part we're most careful about. We ran the teacher model twice over the same 40 inputs at temperature 0. It agreed with itself 82% of the time. Severity is a judgement call, not a fact — so 74% is about 90% of the achievable ceiling, not 74% of a perfect score. We would rather explain that than quote a number that flatters us.
The zero-hallucination result matters more than it looks. Every identifier in the output is checked against the identifiers present in the input. Zero inventions is what makes it safe for Fixer to generate remediation from these reports — a report that references a column which doesn't exist would produce a fix against nothing.
Making it adaptable
Rules get things wrong, and a system that can't be corrected is a system nobody deploys. Two mechanisms, both stored in DataHub rather than in code:
Exceptions. A data owner can suppress a finding ("this replica is deliberate, another team tracks it") or force one the rules missed ("this free-text column contains customer names"). Scoped to a whole asset, a single column, or a specific duplicate pair. A reason is mandatory — an unexplained suppression is a liability six months later.
Configuration. PII patterns, duplicate thresholds, tier rules and
freshness defaults all live in DataHub. Built-in defaults mean it works with
zero setup; a company overrides only what they care about. Adding acct_no
to the bank-account patterns is one CLI call, not a code change.
Challenges we ran into
DataHub's aspect model is stricter than it looks. We spent well over an
hour discovering that corpGroup doesn't support institutionalMemory at
all, and that corpGroupInfo silently drops customProperties — the write
returns 200 and the data simply isn't there. We eventually stored config on a
dataset entity instead. This is undocumented, and it's what we want to
contribute back.
Writes are asynchronous. A successful ingestProposal means the proposal
was accepted, not applied. We chased phantom failures for a while before
realising we were reading back before Kafka had processed the write. Several
"bugs" were just impatience.
One incident per asset was wrong. Our first write-back stored the
highest-severity finding per asset. But customer_features has four problems
at once — stale, unowned, untagged PII, broken upstream. Storing one lost
three, and Fixer could only ever fix one of them. Moving to an array meant
patching every component that read it: Fixer, Historian, and the API each
assumed a single object.
Fixer was inventing owners. Our first version guessed a team name from the table name and assigned it. That's worse than leaving the gap visible — now there's a wrong name attached and nobody notices the problem. We changed it to propose an owner inferred from upstream lineage, with the evidence written out, and let a human confirm.
Compute. We started on AMD MI300X and ran out of credits mid-project, so we moved to an AWS g5.xlarge. The 16GB of system RAM then became the constraint — merging the LoRA into the base model in fp16 OOM-killed the box twice. We ended up skipping the merge entirely and letting llama.cpp apply the adapter at load time.
What we learned
Deterministic where correctness matters, model where judgement does. This was the design decision everything else followed from. Detection has to be exact — a regex match is a fact. Severity is a judgement that needs weighing blast radius against asset tier against exposure. Putting the model on the wrong side of that line would have made the system unreliable in the way that matters most.
Verification is the honest part. It would have been easy to have Fixer report its own success. Having Historian independently re-run the detector is what turns "we applied a fix" into "the problem is gone." It also means our knowledge base has rows that say not yet measured rather than 0% — because Fixer only proposed those, and nothing was verified. Showing them as 0% would misrepresent the system.
A teacher's self-consistency is the real ceiling. We nearly retrained on a 74% severity match before thinking to check what the teacher scored against itself. Knowing the ceiling changed the conclusion completely.
What's next
- Contributing back to DataHub — documentation of the entity/aspect compatibility gaps we hit, which cost us hours and aren't written down anywhere.
- A second adapter for code generation. Fixer currently templates its proposals. Training on (incident + code file → patch) pairs would let it generate real migrations.
- Dismissal learning. Exceptions are set manually today. Historian already tracks outcomes; extending it to track rejected findings would let Scout suppress patterns humans keep dismissing.
- Real-catalog validation. Everything is tested against DataHub's sample catalog plus a fintech ML pipeline we seeded. Running it against a production catalog is the honest next test.
Built With
- amazon-web-services
- datahub
- docker
- fastapi
- gguf
- huggingface
- javascript
- llama.cpp
- peft
- python
- qlora
- qwen
- react
- transformers
- vite


Log in or sign up for Devpost to join the conversation.