Chapter 5: Fine-tuning
Learn how to adapt pretrained LLMs to specific tasks and domains. Master the spectrum of fine-tuning approaches from full parameter updates to parameter-efficient methods like LoRA and QLoRA, understand instruction tuning that transforms base models into helpful assistants, and learn best practices for data preparation and evaluation.
Chapter Overview
Pretraining gives an LLM broad language understanding, but a pretrained model is just a next-token predictor---it does not follow instructions, answer questions helpfully, or refuse harmful requests. Fine-tuning bridges this gap by adapting the model to specific behaviors or domains using curated datasets.
There are two major paradigms of fine-tuning. Full fine-tuning updates all model parameters, achieving the highest quality but requiring significant compute (often 10-100x less than pretraining, but still substantial for large models). Parameter-efficient fine-tuning (PEFT) methods like LoRA update only a small fraction of parameters, making it possible to fine-tune 70B+ models on a single GPU.
The most impactful form of fine-tuning for modern LLMs is instruction tuning: training on (instruction, response) pairs to make the model follow user requests. This is what transforms a base model like Llama into a chat model like Llama-Chat. Combined with RLHF (Chapter 6), instruction tuning produces the helpful, harmless assistants we interact with today.
This chapter covers:
- Transfer Learning: The theoretical foundation for why fine-tuning works
- Full Fine-tuning: When and how to update all parameters
- LoRA & QLoRA: Parameter-efficient fine-tuning that democratized LLM adaptation
- Instruction Tuning: Creating helpful assistants from base models
- Data Preparation: Building high-quality fine-tuning datasets
Chapter Roadmap
Click any topic to jump in
Transfer Learning
Why pretrained knowledge transfers to downstream tasks — catastrophic forgetting, feature extraction, and domain adaptation.
Full updates vs parameter-efficient methods
Full Fine-tuning
Updating all parameters with SFT — when maximum quality justifies the compute cost and overfitting risks.
LoRA & QLoRA
Low-rank adaptation that trains <1% of parameters — the method that democratized fine-tuning on consumer GPUs.
Instruction formats and data quality
Instruction Tuning
Transforming base models into helpful assistants with instruction-response pairs and chat templates.
Data Preparation
Quality over quantity — decontamination, formatting consistency, and augmentation for fine-tuning datasets.
Transfer learning is the principle that knowledge learned on one task can be applied to another. For LLMs, the pretrained model has learned language structure, factual knowledge, and reasoning patterns from trillions of tokens. Fine-tuning transfers this knowledge to specific tasks with far less data than training from scratch.
Key insight: Fine-tuning works because language has universal structure. Syntax, semantics, and reasoning patterns learned from web text transfer to medical reports, legal documents, and code.
In this topic
Why Transfer Learning Works for LLMs
During pretraining, the model learns: (1) Language fundamentals (grammar, syntax, coherence) in early layers, (2) Semantic understanding (word relationships, topic modeling) in middle layers, (3) Task-specific patterns (reasoning, QA, summarization) in later layers. Fine-tuning primarily adjusts the later layers while preserving the foundational knowledge in early layers.
Transfer learning works because the pretrained loss landscape has already converged to a region with useful features. Formally, fine-tuning from pretrained weights finds starting from instead of random initialization. The Hessian at has most eigenvalues near zero (the model has learned a low-dimensional manifold), so fine-tuning primarily moves along a small number of important directions. This is why fine-tuning converges in 3-5 epochs while pretraining takes hundreds of passes over the data.
A pretrained model has never seen radiology reports. Why can fine-tuning on 10,000 reports make it good at radiology?
Catastrophic Forgetting
When fine-tuning on a narrow task, the model can 'forget' knowledge from pretraining. The parameters shift to optimize for the fine-tuning distribution, degrading performance on other tasks. Mitigations: lower learning rates, freezing early layers, mixing pretraining data into fine-tuning, or using regularization techniques.
Catastrophic forgetting occurs when fine-tuning gradients are large in directions that are important for pretrained capabilities. The EWC (Elastic Weight Consolidation) regularizer weights the penalty by the Fisher information , which measures how much parameter matters for the pretrained distribution. Parameters with high Fisher information are 'important' for pretraining and penalized more strongly during fine-tuning.
After fine-tuning on legal text, a model can no longer write Python code. What happened?
Feature Extraction vs Fine-Tuning
Feature extraction: Freeze the pretrained model, only train a new output head. Fast but limited---the frozen representations may not capture task-specific nuances. Fine-tuning: Update all (or some) pretrained parameters. More powerful but risks forgetting and requires more compute. For LLMs, fine-tuning dominates because the generation task requires full model adaptation.
Feature extraction uses the pretrained model as a fixed function and only trains a head with parameters. The capacity of the combined model is limited by the representational power of — if the pretrained features don't capture task-specific information, no amount of head training will help. Fine-tuning has trainable parameters, giving strictly more capacity, but the sample complexity scales as where is the number of fine-tuning examples — risking overfitting when is small.
For sentiment classification with 100 examples, would you fine-tune or use feature extraction?
Domain Adaptation
Continued pretraining on domain-specific text before task-specific fine-tuning. For example: general pretrain -> medical literature continued pretraining -> clinical QA fine-tuning. This two-stage approach often outperforms direct fine-tuning because the model first adapts its language model to the domain, then learns the specific task format.
Continued pretraining on domain text reduces the KL divergence between the domain's language distribution and the model's learned distribution. Without domain adaptation, the model must generalize from the pretraining distribution — the error is bounded by where is the domain divergence. Continued pretraining directly reduces this divergence term, tightening the error bound before task-specific fine-tuning even begins.
You have: a general 7B LLM, 50GB of biomedical papers, and 5000 QA pairs for drug interaction queries. What is the optimal fine-tuning pipeline?
Theory Exercise
Problem:
A pretrained LLM achieves 45% on a legal reasoning benchmark. After fine-tuning on 10K legal examples, it reaches 78%. But its general language ability drops from 92% to 71%. How would you maintain both capabilities?
Hints:
- Think about multi-task learning
- Consider parameter-efficient methods
- Think about data mixing strategies