Inspiration

Most students who want to build a language model do not have a GPU budget. We wanted to find out how far one ordinary consumer computer can go under the Track 01 rule, with no cloud and no pretrained weights, while the same computer stays usable for everyday work, and to report every cost number honestly.

What it does

Laptop-50M is a 49,295,872-parameter decoder-only language model (tied embedding and output head counted once, printed by count_params.py). It was trained from scratch on one Apple M3 iMac (16 GB unified memory, PyTorch MPS, bf16) on 294.9M tokens of FineWeb-Edu.

Cost axis Value
Hardware 1 × Apple M3 iMac (8-core CPU, 10-core GPU, 16 GB)
Training compute 6·N·D ≈ 8.7 × 10^16 FLOPs
Full-speed training time ≈ 20.5 h (294.9M tok ÷ 3,995 tok/s median)
Wall-clock 39.6 h (the Mac stayed in daily use; training paced itself)
Peak memory ≈ 4.7 GB
Cloud spend $0
Full evaluation 535 s on the same GPU

Results (lm-evaluation-harness 0.4.13, 0-shot, full evaluation sets)

Task Metric Laptop-50M Chance
HellaSwag (10,042) acc_norm 27.30 25.0
ARC-Easy (2,376) acc 39.48 25.0
PIQA (1,838) acc 56.75 50.0
WinoGrande (1,267) acc 50.67 50.0
WikiText-103 test word perplexity 105.03 —
WikiText-103 test bits per byte 1.256 —
WikiText-103 test token perplexity (sliding window 512/256) 46.05 —

WikiText-103 is out-of-domain (the WikiText dataset was not used in training). Word perplexity and bits per byte do not depend on the tokenizer, so they compare across models. HellaSwag and WinoGrande are near chance at this scale; we report them as measured. Raw JSON: results/eval_l50m-v1.json.

How we built it

  • Spend the budget on depth, not vocabulary. The rule counts embeddings and the output head. At width 512 a GPT-2 vocabulary with an untied head would need 51.5M parameters before any layer. We tied the input/output embedding and trained our own 16,384-token byte-level BPE, so the embedding is 8.4M (17%) and 83% of the budget (40.9M) goes into 12 transformer blocks (RoPE, SwiGLU 1,536, RMSNorm, 8 heads, context 512).
  • Training that shares the computer. Heavy stages start only after ≥ 120 s of user inactivity (idle gate). While the user is active, a duty-cycle pacer holds the GPU at about 25%, and when idle it runs at full speed (~4,000 tok/s). nice 20, checkpoints every 15 min and on SIGTERM; one resume after a reboot at step 8,961.
  • Memory: a first fp32 run at batch 8 × 1024 used 14.3 GB and swapped (101–211 tok/s); bf16 at 4 × 512 brought it to ~4.7 GB and ~4,000 tok/s.
  • Optimizer: AdamW (0.9, 0.95), weight decay 0.1, WSD schedule (300 warmup → 1e-3 → linear decay to 1e-4 over the last 20%), 9,000 steps × 32,768 tokens.
  • Evaluation: a custom TemplateLM adapter runs lm-evaluation-harness on our model; tests check the scores match a manual computation and do not change with batch size.
  • Engineering discipline: spec first (SPEC.md, acceptance tests AC-1…AC-15), 23 tests, clean architecture (domain / application / adapters / infrastructure).

Challenges we ran into

  • Fitting a useful depth into 50M when the vocabulary alone could exceed the budget.
  • 16 GB unified memory shared with a desktop in daily use: fp32 swapped, so we moved to bf16 and a smaller micro-batch with gradient accumulation.
  • An unplanned reboot at step 8,990: resumable checkpoints with optimizer state brought training back from step 8,961.
  • Honest metrics: the value logged during training (57.44 token perplexity on 25 random windows) differs from the sliding-window value (48.76 val / 46.05 test). The README explains both.

Accomplishments that we're proud of

  • A complete from-scratch LM on a single consumer computer, at zero cloud cost, with the full cost profile (FLOPs, hours, memory) reported.
  • Validation loss fell from 5.566 (step 250) to 3.410 (step 9,000) over 36 checks; it went up only twice, by at most 0.003.
  • Anyone can reproduce it: three commands for data, training and evaluation, and released weights plus tokenizer. Loading the released weights on the CPU gives ARC-Easy 39.48 and PIQA 56.75 again in 70 s (README "Quick check").

What we learned

Under a parameter cap that counts embeddings, the vocabulary size is the biggest design lever. On a shared machine, pacing the training loop costs wall-clock time (39.6 h vs ~20.5 h) but no extra compute.

What's next

More tokens at the same size (the loss was still falling), and a longer decay phase.

AI tool use (disclosure)

Claude Code (Anthropic Claude) wrote the spec, tests, all source code, the data pipeline, the training and evaluation runs, and drafted the README and this text, working under the entrant's direction. No AI model is part of the trained model: random initialisation, FineWeb-Edu text only, tokenizer trained from scratch, no pretrained weights, no fine-tuning, no distillation, no AI-generated training data.

Team

Solo entry: Taegeol Kim (Changwon National University, Korea).

Built With

  • apple-silicon
  • claude-code
  • fineweb-edu
  • lm-evaluation-harness
  • matplotlib
  • metal
  • mps
  • numpy
  • python
  • pytorch
  • tokenizers
  • wikitext-103
Share this project:

Updates

Submission history