PIXELBANKv9.1.0
Menu
Back to LLM Study Plan
Week 5-6

Chapter 5: Fine-tuning

Learn how to adapt pretrained LLMs to specific tasks and domains. Master the spectrum of fine-tuning approaches from full parameter updates to parameter-efficient methods like LoRA and QLoRA, understand instruction tuning that transforms base models into helpful assistants, and learn best practices for data preparation and evaluation.

Chapter Overview

Pretraining gives an LLM broad language understanding, but a pretrained model is just a next-token predictor---it does not follow instructions, answer questions helpfully, or refuse harmful requests. Fine-tuning bridges this gap by adapting the model to specific behaviors or domains using curated datasets.

There are two major paradigms of fine-tuning. Full fine-tuning updates all model parameters, achieving the highest quality but requiring significant compute (often 10-100x less than pretraining, but still substantial for large models). Parameter-efficient fine-tuning (PEFT) methods like LoRA update only a small fraction of parameters, making it possible to fine-tune 70B+ models on a single GPU.

The most impactful form of fine-tuning for modern LLMs is instruction tuning: training on (instruction, response) pairs to make the model follow user requests. This is what transforms a base model like Llama into a chat model like Llama-Chat. Combined with RLHF (Chapter 6), instruction tuning produces the helpful, harmless assistants we interact with today.

This chapter covers:

  • Transfer Learning: The theoretical foundation for why fine-tuning works
  • Full Fine-tuning: When and how to update all parameters
  • LoRA & QLoRA: Parameter-efficient fine-tuning that democratized LLM adaptation
  • Instruction Tuning: Creating helpful assistants from base models
  • Data Preparation: Building high-quality fine-tuning datasets

Chapter Roadmap

Click any topic to jump in

1
Transfer Learning

Why pretrained knowledge transfers to downstream tasks — catastrophic forgetting, feature extraction, and domain adaptation.

Why Transfer Learning Works for LLMsCatastrophic ForgettingFeature Extraction vs Fine-TuningDomain Adaptation
Two approaches to adapting pretrained models

Full updates vs parameter-efficient methods

2
Full Fine-tuning

Updating all parameters with SFT — when maximum quality justifies the compute cost and overfitting risks.

Full Fine-Tuning SetupSupervised Fine-Tuning (SFT)Hyperparameter SensitivityWhen to Use Full Fine-Tuning
3
LoRA & QLoRA

Low-rank adaptation that trains <1% of parameters — the method that democratized fine-tuning on consumer GPUs.

LoRA: Low-Rank AdaptationWhy Low-Rank WorksQLoRA: Quantized LoRALoRA Hyperparameters
Applying fine-tuning in practice

Instruction formats and data quality

4
Instruction Tuning

Transforming base models into helpful assistants with instruction-response pairs and chat templates.

What is Instruction TuningInstruction Dataset DesignMulti-Task Instruction TuningChat Templates and Formatting
5
Data Preparation

Quality over quantity — decontamination, formatting consistency, and augmentation for fine-tuning datasets.

Data Quality Over QuantityFormatting and ConsistencyDecontaminationData Augmentation

Transfer learning is the principle that knowledge learned on one task can be applied to another. For LLMs, the pretrained model has learned language structure, factual knowledge, and reasoning patterns from trillions of tokens. Fine-tuning transfers this knowledge to specific tasks with far less data than training from scratch.

Key insight: Fine-tuning works because language has universal structure. Syntax, semantics, and reasoning patterns learned from web text transfer to medical reports, legal documents, and code.

In this topic

1Why Transfer Learning Works for LLMs
2Catastrophic Forgetting
3Feature Extraction vs Fine-Tuning
4Domain Adaptation
1 of 4
Why Transfer Learning Works for LLMs

During pretraining, the model learns: (1) Language fundamentals (grammar, syntax, coherence) in early layers, (2) Semantic understanding (word relationships, topic modeling) in middle layers, (3) Task-specific patterns (reasoning, QA, summarization) in later layers. Fine-tuning primarily adjusts the later layers while preserving the foundational knowledge in early layers.

Mathematical Intuition

Transfer learning works because the pretrained loss landscape has already converged to a region with useful features. Formally, fine-tuning from pretrained weights θ0\theta_0 finds θ∗=arg⁡min⁡θLtask(θ)\theta^* = \arg\min_\theta \mathcal{L}_{\text{task}}(\theta) starting from θ0\theta_0 instead of random initialization. The Hessian H=∇2LH = \nabla^2 \mathcal{L} at θ0\theta_0 has most eigenvalues near zero (the model has learned a low-dimensional manifold), so fine-tuning primarily moves along a small number of important directions. This is why fine-tuning converges in 3-5 epochs while pretraining takes hundreds of passes over the data.

Example:

A pretrained model has never seen radiology reports. Why can fine-tuning on 10,000 reports make it good at radiology?

2 of 4
Catastrophic Forgetting

Ltotal=Ltask+λ∑i(θi−θipretrained)2\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{task}} + \lambda \sum_i (\theta_i - \theta_i^{\text{pretrained}})^2

When fine-tuning on a narrow task, the model can 'forget' knowledge from pretraining. The parameters shift to optimize for the fine-tuning distribution, degrading performance on other tasks. Mitigations: lower learning rates, freezing early layers, mixing pretraining data into fine-tuning, or using regularization techniques.

Mathematical Intuition

Catastrophic forgetting occurs when fine-tuning gradients ∇Ltask\nabla \mathcal{L}_{\text{task}} are large in directions that are important for pretrained capabilities. The EWC (Elastic Weight Consolidation) regularizer λ∑iFi(θi−θipre)2\lambda \sum_i F_i(\theta_i - \theta_i^{\text{pre}})^2 weights the penalty by the Fisher information Fi=E[(∂log⁡P/∂θi)2]F_i = \mathbb{E}[(\partial \log P / \partial \theta_i)^2], which measures how much parameter θi\theta_i matters for the pretrained distribution. Parameters with high Fisher information are 'important' for pretraining and penalized more strongly during fine-tuning.

Example:

After fine-tuning on legal text, a model can no longer write Python code. What happened?

3 of 4
Feature Extraction vs Fine-Tuning

Feature extraction: Freeze the pretrained model, only train a new output head. Fast but limited---the frozen representations may not capture task-specific nuances. Fine-tuning: Update all (or some) pretrained parameters. More powerful but risks forgetting and requires more compute. For LLMs, fine-tuning dominates because the generation task requires full model adaptation.

Mathematical Intuition

Feature extraction uses the pretrained model as a fixed function fθ:X→Rdf_\theta: \mathcal{X} \to \mathbb{R}^d and only trains a head gϕ:Rd→Yg_\phi: \mathbb{R}^d \to \mathcal{Y} with ∣ϕ∣≪∣θ∣|\phi| \ll |\theta| parameters. The capacity of the combined model is limited by the representational power of fθf_\theta — if the pretrained features don't capture task-specific information, no amount of head training will help. Fine-tuning has ∣θ∣+∣ϕ∣|\theta| + |\phi| trainable parameters, giving strictly more capacity, but the sample complexity scales as O(∣θ∣/n)O(|\theta|/n) where nn is the number of fine-tuning examples — risking overfitting when nn is small.

Example:

For sentiment classification with 100 examples, would you fine-tune or use feature extraction?

4 of 4
Domain Adaptation

Continued pretraining on domain-specific text before task-specific fine-tuning. For example: general pretrain -> medical literature continued pretraining -> clinical QA fine-tuning. This two-stage approach often outperforms direct fine-tuning because the model first adapts its language model to the domain, then learns the specific task format.

Mathematical Intuition

Continued pretraining on domain text reduces the KL divergence DKL(Pdomain∥Pθ)D_{\text{KL}}(P_{\text{domain}} \| P_\theta) between the domain's language distribution and the model's learned distribution. Without domain adaptation, the model must generalize from the pretraining distribution PgeneralP_{\text{general}} — the error is bounded by ϵtask≤ϵbest+dH(Pgeneral,Pdomain)\epsilon_{\text{task}} \leq \epsilon_{\text{best}} + d_{\mathcal{H}}(P_{\text{general}}, P_{\text{domain}}) where dHd_{\mathcal{H}} is the domain divergence. Continued pretraining directly reduces this divergence term, tightening the error bound before task-specific fine-tuning even begins.

Example:

You have: a general 7B LLM, 50GB of biomedical papers, and 5000 QA pairs for drug interaction queries. What is the optimal fine-tuning pipeline?

Theory Exercise

Problem:

A pretrained LLM achieves 45% on a legal reasoning benchmark. After fine-tuning on 10K legal examples, it reaches 78%. But its general language ability drops from 92% to 71%. How would you maintain both capabilities?

Hints:
  • Think about multi-task learning
  • Consider parameter-efficient methods
  • Think about data mixing strategies