GPT-2 Parameter Counter

Adjust the shape, vocabulary, and working context of a GPT-2-style transformer and see where every parameter lives.

Model dimensions

Width of each token representation
Repeated attention + MLP blocks
Entries in the tokenizer vocabulary
Maximum sequence positions
Architecture options
Reference sizes
Total parameters 124.4M
124,439,808 exactly
Token embeddingsvocab × dimension 38.6M
Position embeddingscontext × dimension 786K
AttentionQKV + output projections 28.3M
Layer norms ×2scale + bias in every block 36.9K
Feed-forward networks4× expansion + projection 56.7M
Final layer normscale + bias 1.54K
Output headshared with token embeddings 0

Uses tied input/output embeddings, QKV bias, 4× MLP expansion, projection and MLP biases, two layer norms per block, and a final layer norm.