← All courses

AI Basics: LLMs and the Hugging Face Transformers Library

Explain how the Transformer architecture works (attention, encoder/decoder) and the LLM lifecycle. Solve common NLP tasks using `pipeline()` and directly via models and tokenizers. Fine-tune pretrained models using the Trainer API and a full PyTorch training loop.

Log in to start →

Course preview

What you will learn

Explore the full curriculum before you enrol. Lessons unlock after purchase or free enrolment.

01

Introduction and Transformer Models

Topics covered

  • Set up a working environment (Colab or local Python) to use the Transformers library.
  • Understand what NLP and large language models (LLMs) are, and how they relate.
  • Master the `pipeline()` function for solving common tasks without training.
1 lesson10 questions~25 min
02

The Transformers Library in Practice

Topics covered

  • Understand what happens "under the hood" of `pipeline()`: tokenizer → model → postprocessing.
  • Learn to load models and tokenizers via `AutoModel`/`AutoTokenizer` and `from_pretrained()`.
  • Master the stages of tokenization: splitting into tokens, converting to IDs, special tokens.
1 lesson10 questions~25 min
03

Fine-Tuning Models and the Hugging Face Hub

Topics covered

  • Prepare a dataset for training: tokenization, dynamic padding, `DataCollator`.
  • Fine-tune a classification model with the high-level `Trainer` API.
  • Write a full training loop in "plain" PyTorch and speed it up with Accelerate.
1 lesson10 questions~25 min
04

The Datasets and Tokenizers Libraries

Topics covered

  • Load data from any source (CSV, JSON, local and remote files).
  • Master dataset transformations: `map`, `filter`, `sort`, `shuffle`, splitting into subsets.
  • Work with large data via memory mapping and streaming.
1 lesson10 questions~25 min
05

Classic NLP Tasks

Topics covered

  • Fine-tune a token classification (NER) model with correctly aligned labels.
  • Perform domain adaptation of a masked language model (MLM).
  • Fine-tune seq2seq models for translation and summarization, master the BLEU/ROUGE metrics.
1 lesson10 questions~25 min
06

Debugging and Working with the Community

Topics covered

  • Develop a systematic approach to reading Python tracebacks and library errors.
  • Master the methodology for debugging a training pipeline: from data to gradient computation.
  • Be able to create a minimal reproducible example (MRE).
1 lesson10 questions~25 min
07

Gradio Demos and Data Labeling with Argilla

Topics covered

  • Build interactive demos of ML models with Gradio (`Interface` and `Blocks`).
  • Publish demos: temporary `share=True` links and permanent hosting on Hugging Face Spaces.
  • Connect models from the Hub to a demo in a single line.
1 lesson10 questions~25 min
08

Fine-Tuning LLMs and Reasoning Models

Topics covered

  • Understand chat templates and their role in conversational LLMs.
  • Perform supervised fine-tuning (SFT) using `SFTTrainer` from the TRL library.
  • Apply LoRA for parameter-efficient fine-tuning on a consumer GPU.
1 lesson10 questions~25 min
09

Final Exam

Final exam
35 questions~60 min