Adjust the shape, vocabulary, and working context of a GPT-2-style transformer and see where every parameter lives.
Model dimensions
Width of each token representation
Repeated attention + MLP blocks
Entries in the tokenizer vocabulary
Maximum sequence positions
Reference sizes
Total parameters
124.4M
124,439,808 exactly
Token embeddingsvocab × dimension
38.6M
Position embeddingscontext × dimension
786K
AttentionQKV + output projections
28.3M
Layer norms ×2scale + bias in every block
36.9K
Feed-forward networks4× expansion + projection
56.7M
Final layer normscale + bias
1.54K
Output headshared with token embeddings
0
Uses tied input/output embeddings, QKV bias, 4× MLP expansion, projection and MLP biases, two layer norms per block, and a final layer norm.