Inspiration

Inspired by AfterQuery's work, https://www.afterquery.com/blog/on-policy-distillation-gdpval, we wanted to explore whether a smaller model could learn useful reasoning behavior from a larger teacher using on-policy distillation.

What it does

We used on-policy distillation from Nemotron-3-Ultra-550B-A55B into Nemotron-3.5-Lightning-30B-A3B. The student generates its own responses, the teacher scores the tokens it produces, and those scores guide updates to the student's LoRA adapters.

How we built it

We built the training pipeline in Python using PyTorch, Hugging Face Transformers, and with vLLM. It connects student sampling, teacher scoring, adapter updates, checkpointing, and matched comparisons against the original model.

Challenges we ran into

Due to our time constraints, as generating responses was a substantial part of the workload, we had to focus a lot of efforts onto sampling efficiency and batching.

Accomplishments that we're proud of

Our distilled student achieved 67.17% on GPQA Diamond (133/198) on our best-performing seed, compared with a 64.14% baseline (127/198), a 3.03 percentage-point improvement.

Importantly, this gain did not come with meaningful regressions on our control benchmarks. On GSM8K, our distilled student scored 87% vs. 86% for the baseline, while on ARC-Challenge, it scored 92% vs. 93%.

We also preserved inference speed: the baseline achieved 323.74 tok/s, compared with 322.90 tok/s for our distilled model.

What's next for Nemo-Astra-3.5-Lightning-BF16

We want to train for longer on a larger, more carefully reviewed dataset, explore learning rates and adapter capacity, and evaluate more broadly across reasoning, coding, and instruction following.

Built With

Share this project:

Updates

Submission history