We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Inspiration

Frontier LLMs are amazing, but they all assume one thing: a data center. We asked the opposite question - what if you have 50 million parameters and zero internet?

For on-device agents, air-gapped security SOCs, and IoT defenders, you can't call an API. You need a model that is tiny enough to fit in <100MB quantized, but smart enough to understand CVE triage, AES-GCM vs CBC, and buffer overflows.

We realized standard LLM recipes completely collapse at this scale, and that's what we wanted to engineer our way out of.

How We Built It: The 50M Budget War

We built Vortex-50M entirely from first principles in torch.nn - no transformers training loop, no pretrained weights.

At 50M, every parameter is a battle. The biggest enemy is the vocab table:

$$P_{embed} = V \times d$$

For a standard $V=150,000$ vocab at $d=512$, that's:

$$150,000 \times 512 = 76,800,000 \text{ params} = 153.6\% \text{ of our budget}$$

You are over budget before a single transformer layer.

Our solution - Depth over Width:

  1. Custom 16K BPE Tokenizer: We trained a 16,384-token tokenizer from scratch on 120K docs. Cost: $$16,387 \times 512 = 8,390,144 \text{ params (16.8% of budget)}$$ By tying lm_head to embeddings, we saved 83.2% of the budget for actual reasoning.
  2. Architecture - 512d x 18L: This let us go deep (18 layers) instead of wide. Deep models reason better at small scales.
  3. Stability for Small-Scale:
    • Grouped Query Attention (8Q/2KV): 4x smaller KV-cache for edge inference.
    • QK-Norm: We apply RMSNorm on Q and K before attention: $Q' = \text{RMSNorm}(Q), K' = \text{RMSNorm}(K)$. This prevents attention entropy collapse where small model heads turn into one-hot spikes.
    • Zero-init residuals: c_proj and down_proj are zero-initialized so the network is a perfect identity at step 0.

Training Pipeline:

  • Pretrain: Base model VTXAI/vortex-50m-16k on 1 Billion tokens (Chinchilla-optimal for 50M: $20 \times N$ tokens) from Cosmopedia + FineWeb-Edu with cosine decay to 10% LR.
  • Domain SFT: Instruction-tuned on VTXAI/cyber-crypto-balanced-qa with a 70% technical / 30% conversational anchor mix to prevent catastrophic forgetting, plus assistant-only loss masking.

All params verified by src/param_count.py: 49,846,528 / 50,000,000 [PASS]

Challenges We Faced

1. Residual Saturation: Pre-LN at 18 layers caused gradient blowup. We had to implement QK-Norm + careful RMSNorm placement. 2. Vocab vs Depth Tradeoff: Every 1K vocab tokens we added cost us half a layer. The 16K sweet spot took 4 tokenizer retrains. 3. Small-Model Forgetting: Pure cyber data made the model a parrot that lost English. The 30% anchor data and weight decay 0.01 were critical. 4. Honest Evals: We refused hand-rolled evals. Porting everything to EleutherAI's lm-evaluation-harness with length-norm was painful but gave us real numbers: 42.32% avg (PIQA 54.13%, WinoGrande 52.41%).

What We Learned

At <50M, you are not just training a model, you are doing micro-architecture budgeting. You learn that $P_{total} = P_{embed} + L \times (P_{attn} + P_{mlp})$ is not a formula, it's a budget sheet.

We learned that techniques ignored in 7B models (like QK-Norm, tying, GQA) are not optimizations at 50M - they are survival requirements.

And most importantly: a 49M model can actually be a domain specialist if you give it the right data mix and enough depth.

Built With

PyTorch (from scratch), Hugging Face Transformers, Tokenizers, EleutherAI lm-eval-harness, Cosmopedia, FineWeb-Edu

Built With

  • datasets
  • pytoch
  • transformers
Share this project:

Updates

Submission history