1. Modell & Hardware Config
n_layer
n_head
n_embd
block_size (Context)
batch_size (Micro)
grad_accum_steps
2. Training-Setup
VRAM (GB)
Training data (Billion tokens)
Live-Analysis
Modelsize
0 M
Parameter (vocab=50304)
Tokens per iteration
0
Ziel für LLMs: ~0.5M (500,000)
Estimated VRAM-Usage (bfloat16)
0 GB
Scaling law (Chinchilla)
...