Inspiration
We selected the RSNA Knee Abnormality Detection dataset (https://www.kaggle.com/competitions/rsna-knee-abnormality-detection) with the expectation that it would present a computer-vision challenge. The dataset contains 4,407 knee MRI studies and 819,078 DICOM files, with twelve findings to predict for each study, evaluated using a macro-averaged AUC.
After reviewing the labels, we realized that training a twelve-target model on just 58 studies isn't feasible. However, each training study includes a free-text radiology report written by a radiologist who has reviewed the knee and documented their observations. Importantly, the test.csv file does not include a Report column, meaning a report can't be used as an input during inference. Its only potential role is as training supervision.
This insight reframed the entire project. It shifted our focus from an architectural challenge to a label-quality problem. Additionally, we had a unique advantage: the 58 labeled studies come with both reports and ground truth. This allows us to evaluate a label extractor based on its performance instead of simply trusting it. As a result, we could investigate a question that most datasets do not permit us to answer:
What it does
In a knee MRI study, the model produces twelve calibrated probabilities for various conditions: ACL, MCL, medial and lateral meniscus injuries, three osteoarthritis compartments, effusion, synovitis, Baker's cyst, contusion, and fracture.
What’s particularly interesting is how the model learns to make these predictions. A pipeline of four independent readers processes 4,407 multilingual radiology reports to generate training labels for 98.7% of the studies that lack them. Each reader is evaluated against 58 studies labeled by humans, rather than being assumed to be accurate from the start. The image model then trains using this machine-generated supervision.
Our best configuration achieves a gold-58 macro AUC of 0.8817 (combining three seeds by rank mean), an improvement from 0.7011 with our initial labeller. Notably, this improvement is solely attributed to enhanced supervision, not changes in the architecture.
Additionally, every study in the test set is guaranteed to produce an output row; if a study is unreadable, it defaults to a probability of 0.5 instead of generating no file at all.
How we built it
Labelling Mechanisms and Experiment Overview
| Labeller | Mechanism |
|---|---|
| Multilingual Regex | Hand-built patterns with negation scoping |
| Zero-Shot NLI | mDeBERTa-v3 multilingual entailment against per-finding hypotheses |
| LLM-3B | Qwen2.5-3B-Instruct, structured per-finding queries, logit-derived probabilities |
| LLM-7B | Qwen2.5-7B-Instruct, same protocol |
The reports we analyzed are not in English. The coverage of languages is as follows: English 42%, Turkish 19%, Portuguese 17%, Spanish 16%, Italian 8%, Greek 7.3% (in its own script), Dutch 3.5%, German 2.5%, French 0.4%. An extractor limited to English discards over half of the corpus, and the soft outputs are blended into an ensemble teacher.
Image Model Adaptation
We utilized an image model tailored to the specific problem. Labels are assigned per study, but evidence of a torn meniscus can be found in three slices of one series, which constitutes a multiple-instance problem. Each slice is independently encoded by a ResNet-34 and pooled using learned attention (ABMIL), allowing the model to focus on significant signals instead of averaging them into insignificance. Each study is reduced to four slots (plane × fluid-sensitivity) with a presence mask, meaning that if a study lacks a particular plane, it contributes nothing from that slot instead of just a bank of zeros. The loss function used is Binary Cross Entropy (BCE) with prevalence-aware pos_weight, and the folds are stratified based on each study's rarest positive label. This approach ensures that validation Area Under the Curve (AUC) for rare targets remains defined.
Critical Augmentation Decisions
Two augmentation decisions significantly impacted the results more than the architecture itself. A horizontal flip changes 'Medial OA' to 'Lateral OA'. The standard augmentation enabled by every image pipeline would have inadvertently corrupted four of our twelve targets. Depth reversal presents the same issue in disguise for sagittal series, where the depth axis represents the medial-lateral axis. Both of these augmentations were structurally excluded, and any attempt to re-enable them would lead to test failures.
The Experiment
We conducted an experiment involving five models, where only the training labels differed. The encoder, image resolution, slice count, slots, batch size, learning rate, schedule, epochs, seed, and split remained identical. The training lasted approximately 19 hours on a consumer GPU:
| Training Labels | Label Quality (Gold-58 AUC) | Image Model (Gold-58 AUC) |
|---|---|---|
| Multilingual Regex | 0.7045 | 0.7011 |
| Zero-Shot NLI | 0.7457 | 0.6872 |
| LLM-3B | 0.7911 | 0.7718 |
| Ensemble-3 | 0.8222 | 0.8004 |
| Ensemble-4 | 0.8495 | 0.8159 |
Key Findings
Approximately 80% of the improvement in the quality of the teacher model is reflected in the student model's performance, with validation loss consistently decreasing alongside it. This indicates that better labels are not only ranked higher but are also more reliable.
In a subsequent round, a hybrid teacher achieved a label AUC of 0.9031. Coupled with laterality canonicalization (where every study is geometrically flipped to maintain consistent handedness, ensuring "medial" conveys the same meaning across samples), the image model reached an AUC of 0.8756 on single-seed tests and 0.8817 across three seeds.
Challenges we ran into
247 GB on a machine with 15.9 GB of RAM. The dataset is a single 246.8 GB archive that expands to 530.6 GB, which does not fit on the disk. Compounding the issue, the DICOM paths are structured as three 64-character UIDs deep, exceeding Windows' 260-character MAX_PATH. This limitation caused Kaggle's own extraction process to fail with a truncated-path FileNotFoundError. Therefore, we never extract the files; instead, the preprocessing reads DICOM files directly from the ZIP, where each file is accessed exactly once, decoding into uint8 volumes that are memory-mapped for training. Given the connection reset during the download, resumption via HTTP Range was essential.
VRAM equals system RAM (15.9 GB each). Any operation that stages data in host memory leads to thrashing. Thus, volumes are transferred across the bus as uint8 and normalized on-device, while augmentation runs on the GPU in batches. Device resolution raises rather than defaults to the CPU because we wanted to avoid the worst-case scenario of a CPU-only PyTorch build silently training for nine hours in DRAM.
Using AMD on Windows, without CUDA. We had to solve three issues before running a single epoch. Pip typically prefers the higher version number from PyPI (2.13.0+cpu) over the ROCm build unless the version is pinned exactly. There was no matching torchvision build for this ROCm wheel, so we redefined ResNet in pure PyTorch using torchvision-compatible parameter names — a solution that also worked for Kaggle, as internet access is disabled there. Additionally, MIOpen's batch normalization kernels are JIT-compiled at runtime against C++ headers that the Windows wheel does not include, causing every call to fail with HIPRTC_ERROR_COMPILATION: 'type_traits' file not found. We wrote a probe that identified the broken operations — convolutions worked fine, but batch normalization and instance normalization did not. As a workaround, we disabled MIOpen for these operations while retaining the fast convolution kernels.
Two bugs exhibiting the same symptom. We encountered exit 1 seconds after launch, with no traceback, occurring twice and stemming from unrelated causes: our PowerShell launcher piped output through Tee-Object | Where-Object, and the Composable Kernel emitted approximately 80,000 debug lines per epoch, overwhelming that pipeline and terminating the child process. After addressing that issue, we still faced problems where DataLoader workers did not survive the detached invocation. Training now starts with a simple python command.
The first run appears to hang, but it doesn’t. MIOpen’s exhaustive kernel search can take 10–30 minutes for a cold tensor shape: epoch 1 took 415 seconds, while epoch 2 took only 13 seconds. We lost valuable time over this before we realized what was happening. It’s essential not to benchmark a cold run.
Verifying that the submission measures what we intend. Kaggle processes a directory while we read from a ZIP file. The differing code paths for the same input can result in a submission scoring a different distribution than the model trained on. To address this, we rebuilt the visible test studies both ways and compared the results: 3 out of 3 studies matched exactly on 8,388,608 voxels each. We also discovered that Kaggle's GPU did not have a compatible kernel image for the preinstalled PyTorch, which meant that inference fell back to the CPU instead of failing outright.
Accomplishments that we're proud of
We answered a question, not just built a model. The relationship between label quality and model quality is examined through a controlled experiment involving five models with one variable, rather than relying on anecdotal evidence. This finding is applicable to any dataset that is scarce in labels.
We achieved a significant result using just one desktop GPU. We increased the gold-58 macro AUC from 0.7011 to 0.8817 on an AMD card without CUDA, running on Windows, using a 247 GB archive that we never unpacked, with a peak VRAM usage of 11.5 GB.
We maintained our integrity, even if it cost us potential gains. We established baseline noise levels: ±0.04 on gold-58 and 0.0071 on the pseudo-holdout. This is why ResNet-50 and per-target attention are categorized as "inside noise" instead of additional improvements. Cross-validation also led us to abandon a tempting idea: although per-target labeler selection seemed to provide a clear benefit in-sample, it lost effectiveness out-of-fold. We stand firmly behind publishing what didn’t work as a crucial part of the project.
Correctness is tested, not assumed. We ensured bit-exact preprocessing parity with the Kaggle code path, implemented tests that would fail if augmentation could potentially swap medial for lateral, conducted probes to verify which GPU operations were effective, maintained a per-experiment provenance freeze (including label hashes, cache digests, code hashes, and split hashes), and guaranteed that submissions still produce a complete file even when a study is unreadable.
What we learned
Evaluate a teacher by its worst target, not its average. The NLI labeller has the second-best macro label AUC but produced the worst image model. Its performance is unbalanced—Medial Meniscus at 0.844 and Effusion at 0.835, contrasted with MCL at 0.505 and Lateral OA at 0.547. A macro average masks the reality that two of its twelve targets are nearly random outcomes. A student cannot learn from a random outcome. This is why ensembles are more effective: they have no "dead" targets.
Combine teachers; don’t fine-tune them. Selecting labellers on a per-target basis seemed advantageous in-sample. Under 20 repeats of 5-fold cross-validation, every selection strategy performed worse than the untuned ensemble (best-1 per target: 0.8217 vs. 0.8495). With 58 validation studies, the observed differences were due to selection noise.
The residual errors reflect disagreements about conventions, not parsing failures. Of the remaining 132 label errors, only 18% were cases where all four labellers agreed and were incorrect—82% were split votes, which aggregation can still resolve. Moreover, reports using terms like "mild" or "trace" are 3.84 times overrepresented among error studies, while negation cues are not (1.07 times). Our readers and the ground-truth annotators draw the positive/negative line at different severities. More regex would not have corrected that.
On a label-scarce dataset, the data serves as the model. Using 384×384 input was worse (0.7874). ResNet-50: internal noise. Per-target attention: internal noise. Every architectural and resolution change affected the metric by only ±0.04, while label changes shifted the metric by 0.05–0.07.
One bug taught us more than any hyperparameter. Our negation cues lacked word boundaries. Without \b, the Spanish term sinovitis contains sin ("without"), and the English words noted and normal contain no, which led to findings negating themselves. That single missing character class resulted in Synovitis scoring an F1 of 0.000 and the entire extractor at 0.440. Fixing this issue improved the score to 0.617.
What's next for Reading the Report, Then the Knee
The error analysis leads us to the next step: a severity-threshold pass on the teacher, specifically targeting the 3.84× "mild/trace" cue. This is where our labeling convention diverges from what our readers expect.
Next, we will implement a millimeter-calibrated 120 mm crop. We've observed a 1.20× physical scale variance across studies that the current pipeline overlooks, as resizing to 256 pixels discards PixelSpacing. This means that the same anatomy is presented to the network at different sizes. A 120 mm crop will fit 100% of studies, raising the effective resolution by 1.33× while eliminating the variance.
Following that, we will utilize a domain-pretrained encoder (RadImageNet / DINOv2). This is a capacity change whose mechanism goes beyond simply having "more parameters."
We are also clear about our limitations: we do not yet have a leaderboard score. While the pipeline is verified from end to end and the Kaggle notebook completes successfully, the final submission API call encounters a 403 Permission 'kernelSessions.get' was denied error. This is due to a token-scope limit, not a problem with the pipeline itself. Lastly, it is important to note that this is a research artifact and should not be used for clinical decision-making.
Built With
- amd-gpu
- attention
- dicom
- hip
- hugging-face
- kaggle
- kagglehub
- llm-pseudo-labeling
- mdeberta
- medical-imaging
- multiple-instance-learning
- numpy
- pandas
- powershell
- pydicom
- pylibjpeg
- pytest
- python
- pytorch
- qwen2.5
- resnet
- rocm
- scikit-learn
- transformers
Log in or sign up for Devpost to join the conversation.