← All courses

Smol Course: adapting small language models to your tasks

Run supervised fine-tuning (SFT) of small models with correct chat templates. Apply LoRA/PEFT for cost-efficient fine-tuning. Align models with human preferences using DPO/ORPO.

Log in to start →

Course preview

What you will learn

Explore the full curriculum before you enrol. Lessons unlock after purchase or free enrolment.

01

Instruction Tuning

Topics covered

  • Understand why a base (pretrained) model is turned into an instruct model.
  • Master chat templates: roles, special tokens, the `apply_chat_template` method.
  • Run supervised fine-tuning (SFT) of a small model with `SFTTrainer` from TRL.
1 lesson10 questions~25 min
02

Cost-Efficient Fine-Tuning: LoRA and PEFT

Topics covered

  • Understand the idea of parameter-efficient fine-tuning (PEFT) and how it differs from full fine-tuning.
  • Understand how LoRA works: low-rank matrices, rank, alpha, target modules.
  • Learn to train LoRA adapters via `peft` + `SFTTrainer`, and to load and merge adapters.
1 lesson10 questions~25 min
03

Preference Alignment: DPO and ORPO

Topics covered

  • Understand why a preference alignment stage is needed after SFT.
  • Understand DPO: preference datasets, implicit reward, the beta parameter.
  • Get familiar with ORPO as a single-stage alternative to SFT+DPO.
1 lesson10 questions~25 min
04

Model Evaluation

Topics covered

  • Understand the role of evaluation in the model development loop and the limitations of benchmarks.
  • Get familiar with standard automatic benchmarks (MMLU, TruthfulQA, and others).
  • Master lighteval for running evaluations and analyzing results.
1 lesson10 questions~25 min
05

Multimodal Models

Topics covered

  • Understand the architecture of vision-language models: a visual encoder + a language model.
  • Learn to use ready-made VLMs (SmolVLM) for image captioning, VQA, document understanding, and video.
  • Master VLM fine-tuning: SFT with multimodal chat templates, LoRA/quantization.
1 lesson10 questions~25 min
06

Synthetic Datasets

Topics covered

  • Understand why and when to generate training data with an LLM.
  • Master generating instruction datasets: prompting, SelfInstruct, EvolInstruct, Magpie.
  • Learn to build preference datasets for DPO/ORPO (the UltraFeedback approach).
1 lesson10 questions~25 min
07

Inference and Agents

Topics covered

  • Master simple inference via `pipeline` from transformers and generation parameters.
  • Understand what production serving (TGI) provides: continuous batching, optimizations, monitoring.
  • Learn to build agents with smolagents: code agents, RAG/retrieval agents, custom tools.
1 lesson10 questions~25 min
08

Final Exam

Final exam
35 questions~60 min