Inspiration
Large language models are usually associated with billions of parameters, massive GPU clusters, and enormous training budgets. I wanted to explore a different question:
How capable can a language model become when it is designed, trained, and evaluated entirely from scratch on consumer hardware?
Snowman-LLM is my attempt to answer that question.
What it does
Snowman-LLM is a compact decoder-only Transformer language model with 49,239,552 trainable parameters, keeping it below the 50M-parameter Track 01 limit.
The submitted model was trained from random initialization — without pretrained weights, fine-tuning an existing LLM, knowledge distillation, or hosted inference APIs.
The final training lineage processed approximately 1.24 billion tokens.
How I built it
Snowman went through several architecture and training iterations before reaching the submitted V3 architecture.
The final model uses:
- 13 Transformer layers
- 512-dimensional hidden states
- 8 attention heads
- Rotary Positional Embeddings (RoPE)
- RMSNorm
- SwiGLU feed-forward networks
- PyTorch scaled dot-product attention
- Tied input/output embeddings
- 10,240-token custom BPE tokenizer
- 512-token training context
The entire model contains 49,239,552 trainable parameters.
Training and experimentation were performed locally on a MacBook Pro with an Apple M1 Pro and 16 GB unified memory, using PyTorch's MPS backend.
Training
Instead of performing one monolithic training run, I developed Snowman through successive scratch-training and continuation stages.
The training corpus included mixtures of educational web text, general text, code, stories, instruction-style data, science/reasoning material, and other publicly available datasets.
The final specialization stage additionally used official training splits from PIQA, ARC-Easy and HellaSwag. Held-out evaluation material was excluded from training, with contamination filtering used to help protect evaluation integrity.
The submitted checkpoint has seen approximately 1.241B cumulative training tokens.
Results
The final submitted checkpoint achieved:
| Benchmark | Result |
|---|---|
| PIQA (acc_norm) | 62.08% |
| ARC-Easy (acc_norm) | 42.21% |
| HellaSwag (acc_norm) | 28.03% |
| WinoGrande (accuracy) | 50.67% |
| WikiText-103 | 46.08 perplexity |
An earlier general-language checkpoint achieved 43.72 WikiText-103 perplexity before the final reasoning-oriented specialization stage, demonstrating the trade-off between general language modeling and task specialization.
Challenges
The biggest constraint was training a language model from scratch with only 16 GB of unified memory.
This required keeping the architecture under 50M parameters, carefully controlling context length and batch sizes, optimizing the training pipeline for Apple Silicon, building and tokenizing large datasets locally, and repeatedly measuring whether additional training was actually improving the model.
Another challenge was evaluation integrity. Benchmark evaluation data was kept separate from training, and contamination filtering was used when constructing later training corpora.
What I learned
Snowman-LLM became an experiment not only in language modeling, but in model architecture, tokenization, data composition, optimization, evaluation, checkpoint selection, and the trade-offs involved in training under strict compute constraints.
The project demonstrates that the complete foundational-model pipeline — from tokenizer and architecture to billion-token training and standardized evaluation — can be explored on consumer hardware without starting from a pretrained model.
Log in or sign up for Devpost to join the conversation.