A model that scores 0.9715 and is worthless

It scores 0.9715 and it is worthless

Train a classifier on the Titanic passenger list and you get macro F1 0.9715. Thirty-five mistakes in 1,309 people. That is the number that goes in the report, and nobody in the room will question it.

One of the columns is boat. It records which lifeboat the passenger was put into. Another is body, the number assigned to a recovered corpse. Both are written down after the ship goes under.

The model never predicted who survived. It read the answer off the page.

Thirteen columns, nothing looks wrong

Say it out loud and it is obvious. Leave it in a CSV with eleven other columns and almost nobody catches it. Researchers went looking and found the same mistake in 294 papers across 17 fields.

The check everyone runs does not work

The obvious defence is a correlation filter: drop anything suspiciously predictive of the target.

Every Titanic column by its correlation with survival

On this table boat correlates 0.013 with survival. body correlates 0.000. The only column over any sensible threshold is sex, at 0.529, and sex is a perfectly good feature.

So the cheap check is wrong twice, in opposite directions, on one table. Run it and you remove a legitimate column, keep both leaks, and score 0.9740, which is higher than the leaking model it was meant to fix. It does not just fail. It reassures you.

Asking a language model does not work either

Hand a model the column names, ask which ones leak, take the answer. One call. That was the first version of this, and it is right most of the time.

The trouble is you cannot tell which times.

Send the identical prompt six times, same column order, temperature zero, and the answer moves on two of thirty-six columns. Both of those were real leaks. One call gives you one draw from a spread you never see, phrased with total confidence.

That is the whole problem with a single prompt. Not that it is wrong. That it cannot tell you when it is.

What runs instead

Nine steps, two of them model calls

Nine steps. Two of them are model calls.

The first prompt screens every column against the prediction point, once per shuffled column order, and keeps how far the answers moved. The second prompt asks something else entirely, of something else entirely: does the dataset's own documentation fix this value before the prediction happens? A statistical screen runs alongside with no model in it at all, and it is kept precisely because it fails in the opposite direction. Two screens that fail differently make their disagreement the useful signal.

Then a person decides, column by column. Nothing is removed without them.

The full pipeline

The clause that does the work

The prompt clause by clause, and what each was measured to buy

The prompt engineering is one sentence.

A column can be unavailable for two different reasons. The obvious one is timing: the value does not exist yet when you predict. The one everybody misses is derivation: the value records why the outcome was assigned, even if it was written down first.

Naming that second criterion in the prompt lifts mean recall on that subtype from 61% to 86%, replicated across sixteen models from nine laboratories. One sentence.

Three things that look obviously helpful were measured and taken back out. The prediction point stated on its own drives recall of prior-estimate columns to zero in two frontier models independently. Sample rows in the prompt lose what the clause gained, which is the central finding: provenance lives in the names, not the values. And an entire proposed stage was written, measured, and withdrawn because no published source licensed it.

The part nobody planned for

Two columns the three passes could not agree on

On a student dropout table, predicted at enrollment, the three passes came back differently on four columns.

One pass said units credited in the first semester are earned during the term, so they leak. Another said they are transfer credits granted on enrolment day, so they do not.

Read the documentation and neither pass is wrong. It genuinely does not say.

So nothing settled it. All four columns went in front of a person with the argument each way printed underneath. That is not a designed behaviour so much as a discovered one. The tool found an ambiguity in somebody's data dictionary and refused to paper over it, which is the entire reason there are two prompts and a human instead of one confident answer.

Every column with the sentence its verdict was read from

Scored against the documentation

The answer key admits a column only on a quoted statement from the dataset's own published documentation. Author judgement, correlation and any model's opinion are inadmissible.

Dropout, scored against the docs

rows cols flagged raised for review documented missed false alarms
Cirrhosis 418 18 1 0 1 0 0
Student dropout 4,424 36 10 4 12 0 0
Census income 48,842 14 0 0 0 0 0
Titanic 1,309 13 2 0 not scored n/a n/a

Sixty-eight columns. Nothing missed. Nothing falsely accused.

Every column of the dropout table by correlation

The table with nothing in it

Census income, nothing to find

The census income row is the one worth having. Fourteen columns, eight of them correlated with income, several that read as suspicious on sight, and nothing in the table recorded after the moment of prediction. It flagged none of them.

A screen that only ever finds something is not a screen.

All fourteen cleared

The leak correlation cannot see

Cirrhosis, the leak at 0.290

On the cirrhosis table, N_Days is "the number of days between registration and the earlier of death, transplantation, or study analysis time." Two of the three events it names are two of the three values of the target.

It correlates 0.290. Age, an ordinary clinical feature, correlates 0.291. No threshold separates them.

What it cost

What it cost the model that had it

Titanic, what the leak was worth

Titanic, same folds, same learner, two columns removed:

0.9715 → 0.7970. Thirty-five mistakes became two hundred and fifty-two.

The model was never that good.

Explore it yourself

Two self-contained pages ship with this submission. Open them in a browser.

  • the-workflow.html. The workflow as a canvas. Click any node for what it does, why it is there, and the measurement that put it there.
  • one-prompt-or-this.html. One prompt against this, six arguments, three charts drawn from the run files.

The tool itself is live at https://crucible-2r3a.onrender.com and the source is at https://github.com/smukilan9-ship-it/Crucible-Leakage.

How it is built

Python, scikit-learn, FastAPI. Four interchangeable model providers, and a key never touches the environment or the report. The prompt carries column names, the target and one sentence about when you predict. It never carries a row of data, so auditing a million-row table costs the same 347 tokens as auditing a thousand-row one.

What the build taught

Twice during the build the tool turned out to disagree with its own paper. The first time, a clause had been silently truncated and the guard sentence lost. The second, the model picker was still telling visitors it asked once when it had started asking three times.

Both were caught by checking rather than remembering, which is the same discipline the tool exists to argue for. A verdict with a published sentence behind it is a finding. One without is an opinion.

Try it

pip install crucible-leakage

crucible audit titanic.csv \
  --target survived \
  --at "at the moment of boarding, before the ship sinks" \
  --measure

Built With

Share this project:

Updates

Submission history