Skip to main content

a free course · high-school math only

Start from zero.
Finish by writing a GPT.

Andrej Karpathy fit the complete algorithm behind ChatGPT into 200 lines of plain Python. This course takes you from your first line of code to writing every one of them yourself, and understanding why each one is there.

Every lesson, the same loop

Watch

A short film builds the intuition before any formula.

Explore

Drag, click, roll the dice: play with the idea live.

Build

Write the real code in your browser. Tests tell you when it’s right.

Check

Predict-then-answer questions that make it stick.

Stuck? A tutor that knows the lesson (and your code) gives hints, never answers. Each lesson then lights up the lines of microgpt you now own.

The road

12 modules, about 51 lessons of 20–30 minutes. Part A builds the thinking tools. Part B climbs Karpathy’s own ladder from a table of counts to a Transformer.

  1. M0The destination
  2. M1Computational thinking & Python
  3. M2Data structures
  4. M3Algorithms
  5. M4Math toolkit (visual first)
  6. M5Language as probabilitytrain0.py
  7. M6Learning = walking downhilltrain1.py
  8. M7Autograd = automatic chain ruletrain2.py
  9. M8Attentiontrain3.py
  10. M9The full Transformertrain4.py
  11. M10Training like a protrain5.py
  12. M11Capstone

Open now

Lesson 0.1 · ~15 min
Train a GPT in your browser
Run all 200 lines, watch the loss fall, meet the names it invents.
Lesson 1.1 · ~25 min
Thinking like a computer
Decomposition, patterns, abstraction, algorithms as recipes.
Lesson 1.2 · ~25 min
Values and variables
Numbers, strings, names for things.
Lesson 1.3 · ~25 min
Decisions and loops
`if`, `for`, `while`.
Lesson 1.4 · ~25 min
Functions
Packaging an idea so you can reuse it.
Lesson 1.5 · ~25 min
Reading a dataset
Load 32,033 names from a file.
Lesson 2.1 · ~25 min
Lists
Ordered collections, indexing, slicing.
Lesson 2.2 · ~25 min
Dicts and sets
Look things up by name; keep only unique items.
Lesson 2.3 · ~25 min
Lists of lists = matrices
Grids of numbers, rows and columns.
Lesson 2.4 · ~25 min
Classes and objects
Bundle data with the operations on it.
Lesson 2.5 · ~25 min
Graphs
Nodes and edges: the shape of every computation.
Lesson 3.1 · ~25 min
Recursion
A function that calls itself.
Lesson 3.2 · ~25 min
Depth-first search
Explore a graph all the way down first.
Lesson 3.3 · ~25 min
Topological sort
Socks before shoes: an order that respects every dependency.
Lesson 3.4 · ~25 min
Randomness and weighted dice
Seeds, shuffles and `random.choices`.
Lesson 3.5 · ~25 min
How much work?
Counting steps, and why GPUs exist.
Lesson 4.1 · ~25 min
Functions and graphs
Inputs, outputs, and their pictures.
Lesson 4.2 · ~25 min
Slope by nudging
The derivative is “nudge the input, watch the output.”
Lesson 4.3 · ~25 min
exp and log
Turning multiplying into adding.
Lesson 4.4 · ~25 min
Vectors and the dot product
Similarity as a number.
Lesson 4.5 · ~25 min
Matrix × vector
Many dot products at once.
Lesson 4.6 · ~25 min
Probability distributions
Numbers that add up to 1.
Lesson 5.1 · ~30 min
Counting letter pairs
Your first language model is a table of counts and a weighted die.
Lesson 5.2 · ~25 min
Temperature
Turning creativity up and down.
Lesson 5.3 · ~25 min
Loss = surprise
Scoring a model with negative log-likelihood.
Lesson 5.4 · ~25 min
The number to beat
Why random guessing scores ln 27 ≈ 3.30.
Lesson 6.1 · ~25 min
One knob
A loss curve you can drag.
Lesson 6.2 · ~25 min
Gradient descent
Step downhill; the learning rate is your stride.
Lesson 6.3 · ~25 min
Many knobs
Thousands of parameters, one rule.
Lesson 6.4 · ~25 min
Logits and softmax
From any numbers to probabilities.
Lesson 6.5 · ~25 min
Embeddings
A learnable lookup table per token.
Lesson 7.1 · ~25 min
Computation graphs
Every calculation is a graph.
Lesson 7.2 · ~25 min
The chain rule
Multiplying exchange rates.
Lesson 7.3 · ~25 min
When paths branch
Gradients from different paths add up.
Lesson 7.4 · ~25 min
Building Value
add, mul, pow, log, exp, relu, one at a time.
Lesson 7.5 · ~30 min
backward(): let the graph do calculus
Reverse topological order plus the chain rule, in 14 lines.
Lesson 8.1 · ~25 min
Position embeddings
Where am I in the word?
Lesson 8.2 · ~25 min
Queries, keys, values
A library search between tokens.
Lesson 8.3 · ~25 min
Scaled dot-product attention
Who should I listen to?
Lesson 8.4 · ~25 min
Causal masking & the KV cache
Only look backwards.
Lesson 8.5 · ~25 min
RMSNorm
Keep the numbers in a healthy range.
Lesson 8.6 · ~25 min
Residual connections
A gradient highway.
Lesson 9.1 · ~25 min
Multi-head attention
Several conversations at once.
Lesson 9.2 · ~25 min
The MLP block
Thinking per token.
Lesson 9.3 · ~25 min
Stacking layers
Depth, and the final head.
Lesson 10.1 · ~25 min
Adam
Momentum plus a step size per parameter.
Lesson 10.2 · ~25 min
Learning-rate decay
Big steps first, small steps later.
Lesson 10.3 · ~25 min
Reading loss curves
From 3.3 to 2.28.
Lesson 11.1 · ~180 min
microgpt from a blank file
Every line, written by you.
Lesson 11.2 · ~25 min
Your own dataset
Pokémon, cities, song titles.
Lesson 11.3 · ~25 min
From microgpt to ChatGPT
Tokenizers, GPUs, scale, and post-training.

Built on ideas from Andrej Karpathy (build it from scratch, spelled out) and Andrew Ng (intuition first, then practice). Not affiliated with either.