-
-
This is the final 48.6M-parameter architecture with 12 Co4 blocks, real 6Q/2KV grouped attention, QK-Norm, and 256-token context.
-
DATA-D-v4 combines four curated sources into a certified 2.259B-token stream with normalization, deduplication, and tokenization.
-
Muon handles 90.65% of matrix parameters while AdamW trains embeddings, norms, and vector parameters in the hybrid optimizer.
-
The 2/83/15 WSD schedule cut validation loss from 2.9145 at 1.5B tokens to 2.6305 at 1.9B, making late decay a key result.
-
Post-training tested screening, ranking, joint specialization, SimPO, and capability-gated RFT/RLVR in one evaluation framework.
-
Recovery restored retention, but stronger reasoning-proxy scores did not improve full GIBC performance, so BASE remained final.
-
1. This recap summarizes the final 48.6M model, 1.9B-token run, language-model results, benchmark scores, and BASE selection.
Inspiration
LatticeLM began with a simple constraint that changed the entire shape of the project: the model had to remain below 50 million trainable parameters. Most modern language-model development is discussed in terms of scale, where additional capability is pursued through larger models, larger datasets, and larger compute budgets. I wanted to investigate the opposite regime and determine how much performance could come from better architectural choices, better optimization, cleaner data, and more disciplined evaluation when increasing the parameter count was not available as the default solution.
Working under that limit forced every design choice to compete for the same small budget. Increasing model width meant spending parameters that could no longer be used in the feed-forward layers or embeddings, while attention design affected both representational capacity and training cost. Training decisions also became more consequential because a poorly chosen optimizer or learning-rate schedule could waste a meaningful fraction of the total compute available to the project. This made LatticeLM less about producing a miniature version of a conventional large model and more about treating the entire training system as an optimization problem.
The project therefore developed into an end-to-end investigation of small-model efficiency. Architecture, dataset construction, optimizer choice, learning-rate scheduling, checkpointing, evaluation, and post-training were all treated as parts of the same system rather than as independent components. The goal was not simply to produce a model that fit beneath the parameter limit, but to understand which decisions actually mattered once the model was trained far enough for short-run assumptions to stop dominating the result.
What it does
The final LatticeLM model contains 48,636,168 trainable parameters and was pretrained for 1.9 billion tokens. Its architecture uses twelve causal Co4 blocks with a model width of 552, a feed-forward width of 1728, six query heads, two key-value heads, QK-Norm, a 4096-token vocabulary, a context length of 256 tokens, and untied input and output embeddings. These choices were selected to use nearly the full parameter budget while preserving enough depth, width, attention capacity, and feed-forward capacity to continue benefiting from long training.
Model at a glance
| Measure | Final model |
|---|---|
| Trainable parameters | 48,636,168 |
| Pretraining | 1.9B tokens |
| Architecture | 12-layer causal Co4 |
| Training hardware | 16-core Google Axion CPU, no GPU/TPU |
| Wall time | 185.646 hours / 7.735 days |
| Whole-run throughput | ≈2,842.9 tokens/s |
| Approx. pretraining compute | 0.554 EFLOP |
| WikiText-103 perplexity | 21.9298 |
LatticeLM is also a complete experimental system rather than only a final checkpoint. The repository contains the machinery used to build and certify the training corpus, run long resumable training jobs, benchmark architecture changes, compare optimizer configurations, evaluate checkpoints on language-model and downstream metrics, conduct post-training experiments, and preserve the state needed to reproduce the final decisions. The project was designed so that training results could be traced back to specific configurations and evaluation procedures rather than depending on undocumented manual experimentation.
The final model was selected using the same philosophy. Post-training produced checkpoints with much stronger scores on a targeted reasoning proxy, but those checkpoints were not automatically promoted. They were also evaluated for language-model retention and broader downstream performance, while the original pretrained model remained eligible throughout the selection process. This prevented a specialized improvement from being mistaken for a general improvement when the evidence did not support that conclusion.
How I built it
Architecture and parameter allocation
The final architecture was the result of several rounds of controlled experiments rather than a single fixed design chosen at the beginning. Earlier phases compared smaller dense and Co4-based models under matched conditions, which showed enough promise to justify continuing the Co4 direction. As the available data and training budgets increased, I revisited those conclusions instead of assuming that results from short runs would continue to hold at larger scales.
The production configuration eventually converged on twelve Co4 layers, a 552-dimensional residual stream, and a 1728-dimensional feed-forward network. The context length was increased to 256 after measuring truncation behavior on the final data distribution, while the vocabulary remained at 4096 tokens. QK-Norm remained in the production model after a matched ablation supported its inclusion, and untied embeddings were retained because the experiments did not justify giving up their additional capacity in exchange for the parameter savings from weight tying.
The resulting model reached 48,636,168 trainable parameters, which placed it close to the 50 million parameter limit without relying on arbitrary padding or oversized safety margins. The objective was to spend the available parameters where they contributed to the model rather than simply maximizing a headline width or depth value. This balance between architectural capacity and budget discipline became one of the central design principles of the project.
The Co4 adaptation applies the learned receptive-stream transform
$$ \operatorname{Co4}(r,c)=\operatorname{ReLU6}\left(r^2+2r+c\left(1+|r|\right)\right) $$
to query, key, and value contexts before causal attention. Here r is a learned latent receptive stream and c is the token-conditioned context. The language-model adaptation preserves this elementwise MOD law and the learned receptive streams while using causal scaled dot-product attention for autoregressive prediction.
![]()
Building DATA-D-v4
The final training corpus, DATA-D-v4, contains 2,259,629,459 certified usable tokens drawn from FineWeb-Edu, Wikipedia, FineWeb, and FineMath. The dataset pipeline was designed to make corpus construction reproducible and auditable, so the training stream was not treated as an opaque collection of text files. Source mixing, normalization, stable document identification, deduplication, tokenization, manifest generation, and certification were all performed before the corpus was accepted for production training.
This process also exposed practical data-quality problems that would have been easy to miss in a simpler pipeline. Some upstream records reused document identifiers, which meant that relying on those identifiers directly could cause distinct documents to be treated as equivalent. The builder was changed to use collision-safe identities that also accounted for the underlying token content, and regression tests were added so that the same problem could not silently reappear later.
The final production run consumed 1.9 billion tokens from this certified pool. Maintaining a larger 2.26-billion-token corpus provided enough headroom for training without designing the run around an exhausted dataset or excessive repetition. It also made later scaling decisions easier to evaluate because the available data was not the immediate bottleneck.
Optimizer design
The final training recipe used a hybrid Muon and AdamW optimizer rather than applying the same update rule to every parameter in the network. Approximately 44,089,344 parameters, or 90.65% of the model, are two-dimensional attention and feed-forward matrices and were trained with Muon. The remaining 4,546,824 parameters, or 9.35%, including embeddings, normalization parameters, the output head, and other non-matrix parameters, were trained with AdamW.
This split was motivated by parameter geometry rather than by a desire to use multiple optimizers for its own sake. The large matrix parameters make up the overwhelming majority of the model and are the parameter type for which Muon was most relevant, while AdamW remained appropriate for the remaining vectors and embedding tables. The final hybrid configuration was only promoted after matched experiments showed that the model-quality improvement justified the systems cost.
Earlier Muon experiments were not silently discarded when the configuration changed. The initial version failed the throughput criterion and remained documented as a rejected configuration, while the later corrected test was evaluated separately. Preserving those distinctions made it possible to understand why the optimizer decision changed instead of presenting the final choice as though it had always been obvious.
Training and the WSD schedule
The final production run used a peak learning rate of 8e-4 with a warmup-stable-decay schedule consisting of 2% warmup, 83% stable training, and 15% cosine decay. This schedule became one of the most important findings in the project because the model continued to improve substantially during the final portion of training rather than merely consuming the remaining token budget.
At approximately 1.5 billion tokens, the model had reached a validation loss of 2.9145. The final decay phase then continued training to 1.9 billion tokens, where the validation loss fell to 2.630494. The improvement was also visible on an independent language-model evaluation, where the final model reached a WikiText perplexity of 21.9298 and 1.413462 bits per byte.
| Evaluation | Final BASE |
|---|---|
| HellaSwag | 27.10% |
| ARC-Easy | 36.79% |
| PIQA | 56.69% |
| WinoGrande | 50.99% |
| WikiText-103 PPL | 21.9298 |
| WikiText-103 BPB | 1.413462 |
The strongest conclusion from this stage was not simply that longer training helped, but that the optimization regime changed the value of the remaining tokens. The final WSD cooldown produced one of the clearest improvements in the full production run and demonstrated that late-stage scheduling remained important even after more than a billion tokens of training.
![]()
Training efficiency and compute
The complete 1.9-billion-token production run was trained entirely on a Google Cloud c4a-standard-16 VM with 16 Google Axion CPU cores and 64 GB of RAM; no GPU or TPU acceleration was used. The recorded run duration was 7.73525463 days, or 185.646 hours, which corresponds to an average whole-run throughput of approximately 2,842.9 tokens per second. CPU records were consistently around 1,550% process utilization on the 16-core host, equivalent to roughly 15.5 cores being occupied on average, or about 96.9% aggregate core utilization.
For cross-project normalization, the production pretraining compute is estimated with the conventional 6NT approximation:
$$ C\approx 6NT =6(48{,}636{,}168)(1.9\times10^9) \approx5.54\times10^{17}\ \text{FLOPs} =0.554\ \text{EFLOP}. $$
Co4 follows the 6NT approximation closely enough at this configuration for the estimate to be useful, but the figure is intentionally limited to the final pretraining run. It does not include optimizer work, data construction, evaluation, or the separate experimental runs that preceded the production model. Reporting those boundaries matters because the project included substantial architecture, optimizer, and post-training experimentation beyond the final checkpoint.
Single-node CPU training context
After completing the production run, I compared LatticeLM against publicly documented CPU-trained language models because the combination of model size, token count, and hardware appeared unusually constrained. That comparison required a narrow definition rather than a broad “CPU-trained model” label. Historical RNN language models reached nominal parameter counts in the billions on single CPU machines by using sampled, hierarchical, class-dependent, partially connected, or other approximate output computation, while ThirdAI's much larger BOLT2.5B used dynamic sparse activation and multiple high-core-count CPU systems. Those are genuine CPU-training results, but their nominal parameter counts do not represent the same per-token training regime as running LatticeLM's ordinary full-model next-token computation for 1.9 billion tokens on one 16-core node.
A targeted search of published papers, public model cards, and open-source repositories conducted on October 1, 2026 found no larger publicly documented causal language model than LatticeLM's 48,636,168 trainable parameters that combined at least one billion tokens of pretraining from random initialization, a single CPU-only physical node, and conventional full-model training without dynamic sparsity, MoE routing, or sampled, hierarchical, class-based, or partially connected output approximations. Recent CPU-only projects around 50–60 million parameters were also found, but the documented runs were short experiments or used datasets far below the billion-token threshold rather than comparable production pretraining campaigns.
I do not treat this as a certified world record. A private, unpublished, deleted, or poorly indexed run could satisfy the same conditions. The value of the comparison is that the criteria describe the actual production constraint rather than a benchmark chosen after the fact: LatticeLM was initialized from scratch, kept on one 16-core CPU node for the complete 1.9-billion-token run, and did not use conditional sparsity or an approximate large-vocabulary objective to reduce its normal training computation. The full CPU-training comparison, including BlackOut, the One Billion Word Benchmark RNN systems, BOLT2.5B, and recent CPU-only near-misses, is documented separately so that the conclusion can be updated if a qualifying public counterexample is found.
Post-training and recovery
After the pretrained model was complete, I investigated whether targeted post-training could add stronger reasoning behavior without sacrificing the language-model quality accumulated during pretraining. The experimental system evaluated supervised fine-tuning, ranking objectives, joint objectives, replay strategies, SimPO, and additional capability-gated paths. The ranking and joint objectives produced a large increase on the targeted reasoning proxy, moving from approximately 26.86% for BASE to close to 89% for the strongest specialized checkpoints.
That increase was meaningful, but it was accompanied by substantial degradation in language-model retention. Because of that tradeoff, I did not treat the proxy score as sufficient evidence that the model had improved overall. Instead, the strongest specialized checkpoints became the starting point for a separate recovery study designed to determine whether their newly acquired behavior could coexist with the capabilities learned during pretraining.
The main recovery experiment continued the specialist for 10 million tokens using DATA-D causal training. As the recovery progressed, DATA-D loss and WikiText behavior moved back toward the pretrained model while the specialized reasoning signal gradually weakened. I then evaluated blends between the recovered specialist and BASE to search for configurations that preserved a useful portion of both behaviors.
| Checkpoint | Reasoning proxy | GIBC mean | Outcome |
|---|---|---|---|
| BASE | 0.2686 | 0.4289 | Selected |
| 5% recovered blend | 0.2764 | 0.4288 | Near-neutral |
| 17% recovered blend | 0.3672 | 0.4151 | Rejected |
| Mixed objective, 15% | 0.3418 | 0.4154 | Rejected |
A conservative 5% blend was almost indistinguishable from BASE on the broader evaluation, while the more specialized 17% blend preserved a substantially higher reasoning proxy score but reduced the GIBC mean. A second experiment used generated reasoning completions together with 20% DATA-D replay to preserve more of the specialized behavior during continued training; its strongest screened candidate reached a reasoning proxy of 0.3418 while remaining within the predefined retention limits, but its GIBC mean was still 0.4154.
These experiments showed that the targeted reasoning proxy, language-model retention, and the broader downstream suite were measuring different aspects of the model. The specialized behavior was real, and recovery successfully restored much of the lost language-model quality, but the broader evaluation did not support selecting those checkpoints over the pretrained model. The final selector therefore retained the 1.9-billion-token BASE checkpoint because it remained the strongest candidate under the complete evaluation criteria.
![]()
Challenges I ran into
One of the most difficult parts of LatticeLM was making sure that an apparent improvement represented the model rather than the measurement pipeline. At this scale, implementation details that might appear minor can change the outcome of an experiment enough to produce a false conclusion. I therefore added verification and certification steps throughout the project instead of assuming that plausible-looking metrics were trustworthy.
The dataset pipeline required careful handling of document identity, deduplication, manifests, and reproducibility because reused upstream metadata could otherwise contaminate the training stream. Long CPU training runs also required atomic checkpoint writes and exact resume behavior so that partial files or interruptions could not destroy hundreds of millions of tokens of progress. These systems constraints became part of the training architecture because the experimental design depended on being able to resume the same run rather than approximating it after a failure.
Evaluation created a different set of problems. Context and continuation tokenization had to preserve the true causal boundary, and post-training methods could not be scheduled accurately using only nominal logical tokens because different objectives processed very different amounts of model computation. The final controllers therefore tracked processed tokens, logical tokens, wall-clock time, and experiment state separately, which made later post-training and recovery work much more reliable.
The final challenge was deciding how to handle results that were interesting but did not support the desired conclusion. Some specialized checkpoints produced very large gains on the targeted reasoning metric, while the broader evaluation remained worse than BASE. Those results are documented directly rather than being omitted or reframed as successes. The final model selection follows the evaluation criteria that were established for the project instead of selecting whichever checkpoint produced the most impressive isolated number.
Accomplishments that I am proud of
I am most proud that LatticeLM developed into a complete small-model research system rather than stopping at an early architecture experiment or a single favorable checkpoint. The final 48.6-million-parameter model was trained from random initialization for 1.9 billion tokens entirely on a 16-core CPU VM, without GPU or TPU acceleration, while the same project infrastructure supported architecture experiments, optimizer comparisons, checkpoint selection, post-training, recovery, and final evaluation. That combination made the result a systems and experimental-design project as much as a single training run.
The production training pipeline also became robust enough to support long-running CPU experiments with exact resume behavior, atomic state, certified datasets, machine-readable evaluation, and controlled comparison between checkpoints. DATA-D-v4 provides more than 2.25 billion usable tokens, the hybrid optimizer assigns update rules according to parameter geometry, and the WSD schedule produced a substantial late-stage reduction in validation loss after the model had already consumed 1.5 billion tokens.
I am also proud of the discipline of the final evaluation process. The project preserved the pretrained model as a control throughout post-training, required specialized candidates to satisfy retention constraints, and compared them against broader evaluation before promotion. That process allowed the final result to remain grounded in the complete evidence rather than in the most visually impressive metric.
What I learned
LatticeLM reinforced that parameter count is only one component of model efficiency. When the model is deliberately kept small, architectural allocation, optimizer choice, dataset construction, and learning-rate scheduling become much more visible because inefficient decisions cannot be hidden behind a very large capacity budget. The final WSD improvement was particularly important because it demonstrated that the value of additional training depends heavily on how those tokens are optimized.
The project also showed that different evaluation families can move in different directions even when all of them are measuring legitimate properties of the same model. Language-model loss, WikiText performance, a targeted reasoning proxy, and downstream multiple-choice benchmarks did not always agree. A specialized objective could clearly strengthen the behavior it was designed to measure without producing a corresponding improvement on a broader benchmark suite.
That experience changed how I think about model evaluation. I would no longer treat evaluation as a final stage that simply assigns a score to a trained model. The evaluation suite, retention criteria, controls, and selection rules need to be designed alongside the training process because they determine what the project is actually optimizing.
What's next
The next phase of LatticeLM will focus on understanding the gap between improvements in language modeling, specialized reasoning behavior, and downstream task performance. The current four-task GIBC suite provides a useful comparison point, but the experiments suggest that it does not capture every improvement visible in language-model loss or every behavior introduced during post-training. A broader evaluation framework that includes more generative and reasoning-oriented tasks would make it possible to measure those differences more directly.
The post-training experiments also suggest a clear technical direction. The model demonstrated that strong specialized behavior can be learned and that much of the resulting retention damage can later be repaired. A stronger training objective would preserve both properties in the same checkpoint instead of requiring a separate recovery stage or interpolation afterward.
LatticeLM will continue to pursue the same underlying research question that motivated the project from the beginning: how much useful capability can be extracted from a small language model when architecture, data, optimization, and evaluation are all treated as parts of the same engineering problem?
Log in or sign up for Devpost to join the conversation.