Inspiration
Inspired by AfterQuery's work, https://www.afterquery.com/blog/on-policy-distillation-gdpval, we wanted to explore whether a smaller model could learn useful reasoning behavior from a larger teacher using on-policy distillation.
What it does
We used on-policy distillation from Nemotron-3-Ultra-550B-A55B into Nemotron-3.5-Lightning-30B-A3B. The student generates its own responses, the teacher scores the tokens it produces, and those scores guide updates to the student's LoRA adapters.
How we built it
We built the training pipeline in Python using PyTorch, Hugging Face Transformers, and with vLLM. It connects student sampling, teacher scoring, adapter updates, checkpointing, and matched comparisons against the original model.
Challenges we ran into
Due to our time constraints, as generating responses was a substantial part of the workload, we had to focus a lot of efforts onto sampling efficiency and batching.
Accomplishments that we're proud of
Our distilled student achieved 67.17% on GPQA Diamond (133/198) on our best-performing seed, compared with a 64.14% baseline (127/198), a 3.03 percentage-point improvement.
Importantly, this gain did not come with meaningful regressions on our control benchmarks. On GSM8K, our distilled student scored 87% vs. 86% for the baseline, while on ARC-Challenge, it scored 92% vs. 93%.
We also preserved inference speed: the baseline achieved 323.74 tok/s, compared with 322.90 tok/s for our distilled model.
What's next for Nemo-Astra-3.5-Lightning-BF16
We want to train for longer on a larger, more carefully reviewed dataset, explore learning rates and adapter capacity, and evaluate more broadly across reasoning, coding, and instruction following.
Log in or sign up for Devpost to join the conversation.