In 2024, Andrej Karpathy released microgpt, a 240-line GPT implementation. Other versions kept shrinking it while preserving a working model.
Smaller GPT Implementations
microgpt240 linesOriginal by Andrej Karpathy
picogpt64 linesShared as a QR code
femtogpt53 linesEncoded as a 3000-digit prime number
attogpt31 linesPreserved on 52 IBM punch cards
attogpt cuts the original line count by 87.1%. Its 31 lines and 3,117 characters still include autograd, multi-head attention, training, and inference.
By the Numbers
31
Lines of Code
52
Punch Cards
3.1KB
Total Size
What It Actually Does
The implementation includes:
Autograd Engine: Automatic differentiation with backward pass
Training Loop: Adam optimizer with bias correction and learning rate decay
Text Generation: Inference with temperature-controlled sampling
After training on a dataset of names for 1,000 steps, it generates plausible new names:
sample 1: kamon
sample 2: ann
sample 3: karai
sample 4: jaire
sample 5: vialan
sample 6: karia
sample 7: yeran
sample 8: anna
sample 9: areli
sample 10: kaina
The Punch Card Format
I also encoded the program as IBM-style punch cards, using the physical 80-column by 12-row format common in early computing.
Each Base64 character maps to a 12-bit punch pattern. This is a modern encoding, not period-correct Hollerith code, but it is reversible and fits the network on 52 cards.
Explore the Punch Cards
Each card shows the actual punch hole patterns. Black holes represent binary 1s.