← All courses

About the Course: Diffusion Models — Image and Audio Generation

explain how diffusion models work (the forward and reverse processes, the noise schedule, sampling); generate images and audio with the 🤗 Diffusers library (pipelines, models, schedulers); train your own diffusion model from scratch and fine-tune pretrained models on new data;

Log in to start →

Course preview

What you will learn

Explore the full curriculum before you enrol. Lessons unlock after purchase or free enrolment.

01

Introduction to Diffusion Models and the Diffusers Library

Topics covered

  • Understand what generative models are and where diffusion models fit among them.
  • Break down the forward process (adding noise) and the reverse process (step-by-step denoising).
  • Master the basic components of the 🤗 Diffusers library: pipelines, the UNet model, schedulers.
1 lesson10 questions~25 min
02

Training Your Own Model and Fine-Tuning

Topics covered

  • Master the full training loop for a diffusion model trained from scratch on your own data.
  • Understand when training from scratch is impractical and why fine-tuning a pretrained model is an efficient alternative.
  • Learn to fine-tune a pretrained model on a new dataset using Diffusers.
1 lesson10 questions~25 min
03

Guidance and Conditional Generation

Topics covered

  • Understand how guidance (control at sampling time) differs from conditioning (a condition set during training).
  • Learn to steer an unconditional model with an arbitrary loss function, including CLIP guidance from text.
  • Break down the ways to feed a condition into a UNet: extra channels, adding embeddings, cross-attention.
1 lesson10 questions~25 min
04

Stable Diffusion

Topics covered

  • Understand what latent diffusion is and why a VAE is needed.
  • Break down the components of Stable Diffusion: VAE, the CLIP text encoder, a UNet with cross-attention, the scheduler.
  • Master classifier-free guidance (CFG) and the effect of the guidance scale.
1 lesson10 questions~25 min
05

Advanced Techniques: Editing, Audio, New Architectures

Topics covered

  • Understand how distillation speeds up diffusion model sampling.
  • Master the main approaches to image editing, including DDIM inversion.
  • Break down how diffusion is applied to video and audio (spectrograms).
1 lesson10 questions~25 min
06

Final Exam

Final exam
35 questions~60 min