← All courses

LLMs From Scratch: Build a GPT-like Model with Your Own Hands

Explain the full LLM lifecycle: data → pretraining → fine-tuning. Implement tokenization (including BPE), embeddings, and a sliding-window data loader. Code attention mechanisms from scratch: scaled dot-product, causal, and multi-head attention.

Log in to start →

Course preview

What you will learn

Explore the full curriculum before you enrol. Lessons unlock after purchase or free enrolment.

01

What Are Large Language Models

Topics covered

  • Understand what an LLM is and how it differs from classic NLP systems.
  • Break down the stages of a model's lifecycle: pretraining → fine-tuning.
  • Get an overview of the course roadmap: which components we'll build and in what order.
1 lesson10 questions~25 min
02

Working with Text Data

Topics covered

  • Implement a simple regex-based tokenizer and a token→ID vocabulary.
  • Understand the purpose of special tokens and the principle behind Byte Pair Encoding (BPE).
  • Build a dataset and a sliding-window DataLoader for the 'predict the next token' task.
1 lesson10 questions~25 min
03

Attention Mechanisms

Topics covered

  • Understand why the transformer needs attention and what the limitations of architectures without it are.
  • Implement, step by step: simple attention without weights → self-attention with trainable Q/K/V matrices → causal attention → multi-head attention.
  • Understand the scaling of dot products and the role of dropout.
1 lesson10 questions~25 min
04

Building the GPT Architecture from Scratch

Topics covered

  • Assemble the full GPT-2 small model (~124M parameters) from its individual components.
  • Implement LayerNorm, the GELU activation, the feed-forward network, and residual connections.
  • Assemble the transformer block and the full `GPTModel`; count the number of parameters.
1 lesson10 questions~25 min
05

Pretraining on Unlabeled Data

Topics covered

  • Implement the standard quality metrics for a generative model: cross-entropy and perplexity.
  • Write a full pretraining loop with a train/val split and periodic sample generation.
  • Learn decoding strategies: temperature scaling and top-k sampling.
1 lesson10 questions~25 min
06

Fine-Tuning for Text Classification

Topics covered

  • Understand the difference between classification fine-tuning and instruction tuning.
  • Prepare the SMS spam dataset: class balancing, padding, DataLoader.
  • Replace GPT's output head with a classification head and choose which layers to fine-tune.
1 lesson10 questions~25 min
07

Instruction Tuning: Teaching the Model to Follow Instructions

Topics covered

  • Understand why supervised instruction fine-tuning (SFT) is needed and how it differs from classification.
  • Prepare an instruction dataset: a prompt template (Alpaca-style), dynamic padding, loss masking.
  • Fine-tune GPT-2 medium (355M) on 1,100 'instruction → response' pairs.
1 lesson10 questions~25 min
08

Final Exam

Final exam
35 questions~60 min