Muta ADTC 2026 — Progress Update
From Gate 1 submission to our current model
At Gate 1, we submitted a fine-tuned Qwen3.5-0.8B Q4_0 model for Muta. Our optimization at the time strongly favored the ADTC combined objective: accuracy, performance, and efficiency on an 8 GB RAM, CPU-only machine. The submitted model achieved a strong efficiency profile and successfully passed the initial screening.
| Gate 1 metric | Score |
|---|---|
| Accuracy & Quality | 39.20 |
| Performance | 56.00 |
| Efficiency | 90.65 |
| Total | 54.53 |
The judges' feedback, however, exposed the main trade-off we had made: we had optimized too aggressively for speed and efficiency at the expense of reliability. The model could produce convincing-looking answers while making early arithmetic mistakes, mixing units or currencies, giving weak scientific analogies, or propagating an incorrect first step through the rest of the solution.
For an educational product, this is not acceptable. From that point onward, we redefined our priority: accuracy means critical-thinking reliability—the ability to reason correctly, follow instructions, remain internally consistent, correct misconceptions, and teach safely.
Gate 2: choosing a stronger intelligence core
We revisited the stronger model we had already identified during Gate 1: our fine-tuned Muta Tutor Qwen2.5-1.5B Q4_K_M. Although it was larger and slower under the scalar profiler, it was consistently stronger on STEM reasoning.
In our matched evaluation, Qwen2.5-1.5B improved ARC-Easy accuracy from 70.2% to 77.8% compared with the submitted Qwen3.5-0.8B, and performed better on several of the judges' mathematics, science, and explanation tasks.
We then widened the search instead of assuming Qwen2.5 was automatically the final answer. We tested a broader field of small models—including MiniCPM5, LFM2.5, Qwen3/3.5 variants, VibeThinker, Falcon-H1, and OpenReasoning Nemotron—across:
- 500 ARC-Easy questions
- 100 custom STEM prompts
- all Gate 1 judge prompts
- CPU speed and memory measurements
Across that field, Muta Tutor Qwen2.5-1.5B remained the strongest overall development control, with:
| Metric | Muta Tutor Qwen2.5-1.5B |
|---|---|
| ARC-Easy | 77.8% |
| Full STEM passes | 57/100 |
| Core-correct STEM | 80/100 |
| Gate 1 re-test | 47/100 |
| Completed judge answers | 10/10 |
| AVX2 score proxy | 84.1383 |
This became our quality-first model going forward.
Improving accuracy: what we tried
We then asked whether the selected Muta Tutor could be made substantially more accurate through further fine-tuning.
We built a 2.5M-row STEM data warehouse spanning mathematics, physics, chemistry, biology, integrated science, tutoring styles, and WAEC/WASSCE material. Rather than train blindly on all 2.5M rows, we created a balanced 300,350-row training sample to control subject imbalance, repetition, leakage, and low-quality rows.
We ran:
- Eight 20K-row BF16 LoRA hyperparameter pilots
- Three full-data training lineages
- A separate science + multi-turn tutoring fine-tune
- Held-out evaluations on science, practical STEM, tutoring quality, judge prompts, and STEM MC
The key result was consistent: lower validation loss did not automatically produce a better tutor. Several new checkpoints looked better during training but regressed on fresh evaluation, tutoring behavior, or generalization.
Our final matched comparison still favored the incumbent Muta Tutor:
| Test | Incumbent Muta | Best science challenger (P2) |
|---|---|---|
| Held-out science | 83.565% | 83.565% |
| Practical 2,000 | 41.325 | 39.850 |
| Tutor quality | 22/64 | 12/64 |
| Judges | 57/100 | 50/100 |
| STEM MC | 32/50 | 32/50 |
This was an important outcome: our earlier model-selection and fine-tuning work had produced a stronger checkpoint than we initially realized. Further training increasingly showed diminishing returns and regression/catastrophic-forgetting risk.
Decision: retain Muta Tutor Qwen2.5-1.5B as the quality-first model.
Optimization: recovering speed without giving up the model
Once the quality-first model was fixed, we changed the question again:
How much can we compress and accelerate the selected Muta Tutor without destroying the capability that made us choose it?
We explored vocabulary pruning, layer pruning, knowledge distillation, MoE conversion, FFN width pruning, Q4_0 quantization, and quantization-aware training (QAT).
The strongest compression path was:
- vocabulary reduced to 32K
- depth reduced to 26 layers
- FFN width reduced to 7168
- distillation on verified teacher data
- QAT under simulated Q4_0 noise
This produced our current high-efficiency deployment variant:
refine-qat100-Q4_0.gguf
1.05B parameters · ~593 MB · 15.51 tok/s · ~706 MB peak RAM
Compared with the published Muta Tutor, scalar decode speed increased from roughly 5.5 → 15.5 tok/s, while peak RAM fell from roughly 1.1 GB → 0.7 GB.
However, the compressed model still sacrifices some of the tutoring and reasoning quality of the original Muta Tutor. We therefore treat it as a high-efficiency deployment variant, not yet as the new quality winner.
Where we are now
We now have two clearly defined model tracks:
| Track | Model | Purpose |
|---|---|---|
| Quality-first | Muta Tutor Qwen2.5-1.5B Q4_K_M | Best current tutoring/reasoning model |
| Efficiency-first | refine-qat100-Q4_0 | ~593 MB, ~15.5 tok/s scalar CPU deployment variant |
The main lesson from Gate 1 to now is that Muta cannot be optimized by chasing a single metric. Accuracy, tutoring quality, throughput, memory, quantization format, CPU kernels, data quality, and training objective all interact.
Our current direction is therefore deliberate: keep the original Muta Tutor as the intelligence benchmark, and only promote a compressed or further fine-tuned model when it matches that benchmark on fresh reasoning and tutoring evaluations—not merely on training loss or speed.
Log in or sign up for Devpost to join the conversation.