LH-Tech AI
🤗 Hugging Face GitHub Reddit
🤗 HF GitHub r/
Research - April 17, 2026

Project Apex 2 350M

The Rematch. Same size, modern architecture, dramatically smarter.

350 Million
Parameters
20 Billion
Training Tokens
512
Context Length
57
Tokens per Param

Why Apex 2 350M crushes Apex 1 350M

Same parameter count. Same hardware. But with the modern Llama architecture (RoPE, SwiGLU, RMSNorm), 2× the training data from FineWeb-Edu, and Flash Attention 2, Apex 2 350M decisively outperforms its predecessor across every benchmark. This is what happens when you combine smart training with proper scale.

Feature Apex 1 350M (old) Apex 2 350M (New)
Architecture Standard nanoGPT Llama (Modern)
Training Tokens 10 Billion 20 Billion
Token Density 28 Tokens/Param 57 Tokens/Param
Attention Standard SDPA Flash Attention 2
Val Loss 2.8008 ~2.60 (Predicted)
HellaSwag (Acc) 38.56% ~44% (Predicted)
Training Time 8 Days ~8-10 Days

Modern Training Recipe

Draft Training Code

Go back home