← All courses

Neural Networks: Zero to Hero

Explain and manually implement backpropagation. Build character-level language models: from bigrams to an MLP, a WaveNet-like network, and a GPT transformer. Diagnose the training of deep networks: activation and gradient statistics, Batch Normalization, initialization.

Log in to start →

Course preview

What you will learn

Explore the full curriculum before you enrol. Lessons unlock after purchase or free enrolment.

01

Micrograd: Backpropagation from Scratch

Topics covered

  • Understand what a gradient is and why it is the central concept in training neural networks.
  • Implement automatic differentiation (autograd) over scalars: the `Value` class, the computation graph, `backward()`.
  • Build and train a tiny multilayer perceptron (MLP) without any deep learning libraries.
1 lesson10 questions~25 min
02

Makemore 1: The Bigram Language Model

Topics covered

  • Understand the language modeling task: predicting the next character.
  • Build a bigram model two ways: by counting frequencies and via a single-layer neural network.
  • Get comfortable with `torch.Tensor`, indexing, broadcasting, and the loss function — negative log-likelihood (NLL).
1 lesson10 questions~25 min
03

Makemore 2: The MLP Language Model

Topics covered

  • Extend the model's context from one character to several via embeddings and a multilayer perceptron (MLP).
  • Learn core ML methodology: train/dev/test splits, learning rate selection, under-/overfitting, hyperparameters.
  • Learn to train a model with mini-batches and read a training curve.
1 lesson10 questions~25 min
04

Makemore 3-4: Activations, BatchNorm, and Manual Backprop Through Tensors

Topics covered

  • Learn to diagnose the "health" of a deep network: activation, gradient, and weight-update statistics.
  • Understand proper initialization (Kaiming) and the Batch Normalization mechanism, along with its benefits and pitfalls.
  • Walk through backpropagation by hand at the tensor level: cross-entropy, linear layers, tanh, batchnorm, embeddings.
1 lesson10 questions~25 min
05

Makemore 5: Building WaveNet

Topics covered

  • Deepen the network hierarchically: from a "flat" MLP to a tree-shaped architecture in the spirit of WaveNet (DeepMind, 2016).
  • Understand how `torch.nn` works under the hood by implementing your own equivalent modules (`Linear`, `BatchNorm1d`, `Tanh`, `Embedding`, `Flatten`, `Sequential`).
  • Get a feel for the real DL development process: tracking tensor shapes, reading documentation, iterating.
1 lesson10 questions~25 min
06

GPT: A Transformer from Scratch

Topics covered

  • Implement a decoder transformer (GPT) from scratch, based on the "Attention Is All You Need" paper (2017).
  • Understand the self-attention mechanism: queries, keys, values, masking, multi-head.
  • Assemble a complete transformer block: attention + feed-forward, residual connections, LayerNorm, dropout — and train a "Shakespeare-style" text generator.
1 lesson10 questions~25 min
07

The GPT Tokenizer: Byte Pair Encoding

Topics covered

  • Understand what tokenization is and why LLMs work with tokens rather than characters.
  • Implement the Byte Pair Encoding (BPE) algorithm from scratch: vocabulary training, `encode()`, `decode()`.
  • Get to grips with Unicode/UTF-8, the GPT-2/GPT-4 regex patterns, and special tokens.
1 lesson10 questions~25 min
08

Final Exam

Final exam
35 questions~60 min