Inspiration
GIBC's Track 01 caps a language model at 50,000,000 trainable parameters, and the cap counts the token embeddings and the output head. Most small models borrow GPT-2's 50,257-token vocabulary without thinking about it. At a width of 512, that one table costs 25.7 million parameters: 51.5% of the whole budget, spent before a single transformer layer exists. I wanted to see what happens if you treat the vocabulary as the first design decision instead of a default.
What it does
The Embedding Tax is a 49,296,896-parameter GPT trained from scratch, plus a site where you can run it yourself.
- The model. A 16,384-token vocabulary costs 8.4M parameters (16.8% of the cap) instead of 25.7M. With everything else held equal, that buys 13 transformer layers instead of 7.
- It runs in your browser. The trained weights are exported to a 50 MB file and run on your own device with WebGPU or WebAssembly. Nothing you type leaves the page, and there is no server behind it.
- The explorer. A calculator where you move the vocabulary size and width and watch how many layers fit under the cap.
- The research page. The training curve, the benchmark scores next to GPT-2 Small and random chance, and the published samples, all read from the run's own log files.
Scored with lm-evaluation-harness on the final checkpoint:
| Task | This model (49M) | GPT-2 Small (124M) | Chance |
|---|---|---|---|
| ARC-Easy | 40.32 | 39.73 | 25.0 |
| PIQA | 60.28 | 62.08 | 50.0 |
| HellaSwag | 28.71 | 31.38 | 25.0 |
| WinoGrande | 50.99 | 50.67 | 50.0 |
WikiText word perplexity is 65.11 (1.1267 bits per byte).
How we built it
- Model: a GPT with 13 layers, 8 heads, width 512, a 1,024-token context, rotary positions, weight-only RMSNorm, no biases, and the embedding tied to the output head. A script prints the parameter count and checks it against the live model.
- Data: FineWeb-Edu (sample-10BT, ODC-By 1.0), streamed and tokenized with a byte-level BPE tokenizer I trained on the same corpus. 1.9 billion training tokens, with 100 million held out that the model never sees.
- Training: one RTX 4090 in PyTorch with bf16. 14,495 steps, 7.3 hours on the clock, about 5.2 hours of actual compute, roughly 5.6 × 10^17 FLOPs. Final validation loss 3.10.
- Evaluation: the real lm-evaluation-harness, with the model registered as a harness model class, so the scoring code is theirs and not mine.
- Browser: exported to ONNX, quantized to int8, and run with onnxruntime-web. The page checks the file's SHA-256 and runs a fixed self-check prompt before it trusts a backend.
- Site: plain HTML, CSS and JavaScript with no build step, hosted on GitHub Pages.
- AI tools: the code, the tooling and the website were written with Claude Code. The model itself is trained from scratch: no pretrained weights, no fine-tuning, no distillation, and no hosted model behind the demo.
Challenges we ran into
- A bigger batch was 60 times slower. Batch 16 ran at about 107,000 tokens per second. Batch 32 collapsed to about 1,700, so I kept the small batch and used gradient accumulation.
- A silent data leak. Re-running the data script would have written validation documents back into the training set. I moved the split inside the script so it can't happen.
- Heat. The GPU could not run flat out for seven hours, so training ran at about 65% load with cooling pauses. A watchdog I wrote to protect the run once stopped it 500 steps early, and I had to resume it.
- The wrong perplexity. Per-token perplexity flatters a small vocabulary. I had to switch to word perplexity and bits per byte, which compare fairly across tokenizers.
- Quantizing without breaking it. The first int8 export only agreed with the original on 89% of tokens. The final one agrees on 70 of 72.
- Phones. The page built two copies of the model at once and iOS killed the tab for memory. Phones now build one, and peak memory dropped by roughly half.
- It writes fluently and is often wrong. Raw sampling looped and drifted off topic, so I tested more than 20 decoding setups with blind judging to find ones that stay on the prompt.
Accomplishments that we're proud of
- A model 2.5 times smaller than GPT-2 Small that edges it on ARC-Easy (40.32 vs 39.73) and stays within a few points on PIQA and HellaSwag, after about five hours on one consumer GPU. WinoGrande sits at a chance for both.
- Anyone can run the actual trained model in a browser tab, on a phone too, with no server and no account.
- Every measured number on the site is read from the run's log and the harness output. The training log, evaluation output, and samples are in the repository, so the results can be checked.
What we learned
- Under a parameter cap that counts embeddings, vocabulary size is an architecture decision. It determined how deep this model could be.
- Throughput is not monotonic in batch size. Measure before assuming.
- Which metric you report matters as much as the number. A metric that depends on your own tokenizer cannot be compared with anyone else's.
- A small model's knowledge is limited by its size, but how you sample from it changes how good it feels to use.
- Getting a model to run on a phone is mostly a memory problem, not a speed problem.
What's next for Embedding-Tax
- A real vocabulary sweep (4k, 8k, 16k, 32k) at the same parameter budget, to measure the tradeoff instead of arguing it from arithmetic.
- A longer training run. This one stopped at 1.9 billion tokens and the loss was still falling.
- A key-value cache in the browser runtime, so writing gets faster as the text gets longer.
- Publishing the full checkpoint so anyone can re-run the benchmarks.
Log in or sign up for Devpost to join the conversation.