The Rematch. Same size, modern architecture, dramatically smarter.
Same parameter count. Same hardware. But with the modern Llama architecture (RoPE, SwiGLU, RMSNorm), 2× the training data from FineWeb-Edu, and Flash Attention 2, Apex 2 350M decisively outperforms its predecessor across every benchmark. This is what happens when you combine smart training with proper scale.
| Feature | Apex 1 350M (old) | Apex 2 350M (New) |
|---|---|---|
| Architecture | Standard nanoGPT | Llama (Modern) |
| Training Tokens | 10 Billion | 20 Billion |
| Token Density | 28 Tokens/Param | 57 Tokens/Param |
| Attention | Standard SDPA | Flash Attention 2 |
| Val Loss | 2.8008 | ~2.60 (Predicted) |
| HellaSwag (Acc) | 38.56% | ~44% (Predicted) |
| Training Time | 8 Days | ~8-10 Days |