PIXELBANKv9.1.0
Menu
Back to VLM Study Plan
Week 8-9

Chapter 8: Parameter-Efficient Fine-Tuning

Master the techniques that make VLM adaptation practical: from LoRA's elegant low-rank decomposition to QLoRA's memory-efficient quantized training, adapter tuning, and prompt tuning. Understand the mathematical foundations, implementation trade-offs, and practical recipes for fine-tuning VLMs on consumer hardware — then learn how to merge, serve, and switch between multiple task-specific adapters in production.

Chapter Overview

Full fine-tuning of a 7B parameter VLM requires storing four copies of every parameter in GPU memory: the parameters themselves (bf16, 14GB), their gradients (bf16, 14GB), and the Adam optimizer states (fp32 master weights + momentum + variance, 84GB). The total: 112GB per GPU — impossible even on the latest 80GB GPUs without multi-GPU setups. For a 70B model, the numbers are 10x worse. This memory barrier makes full fine-tuning impractical for most research labs and entirely impossible for production teams that need to adapt VLMs to dozens of tasks.

Parameter-Efficient Fine-Tuning (PEFT) methods solve this by training only a small number of additional parameters while keeping the pretrained weights frozen. The insight is profound: the adaptation from a general model to a task-specific one can be captured in a surprisingly low-dimensional subspace. LoRA demonstrates this by decomposing weight updates into low-rank matrices, reducing trainable parameters by 100-1000x while preserving 95%+ of full fine-tuning performance.

The key PEFT methods form a hierarchy of approaches:

  1. LoRA (Low-Rank Adaptation): Adds trainable low-rank matrices B⋅AB \cdot A alongside frozen weights. The most widely used PEFT method, effective and simple.
  2. QLoRA: Combines LoRA with 4-bit quantization of the base model, enabling 65B model fine-tuning on a single 48GB GPU.
  3. Adapter Tuning: Inserts small bottleneck layers between transformer blocks. Pioneered PEFT but largely superseded by LoRA.
  4. Prefix/Prompt Tuning: Prepends learned "soft prompts" to the input, modifying model behavior without changing any weights.

For VLMs specifically, PEFT raises a unique question: which components should be adapted? The vision encoder, the projection layer, and the LLM backbone each contribute differently to task performance. Applying LoRA to the wrong components wastes parameters; applying it to the right ones can match full fine-tuning at 0.1% of the parameter cost.

This chapter covers each PEFT method in depth: the mathematical foundations (why low-rank works, what information is captured in the rank-rr subspace), practical implementation (which layers, what rank, what learning rate), and the production story (merging adapters for zero-overhead inference, serving multiple tasks from a single base model, and combining adapters via model arithmetic).

The formalism for LoRA begins with a simple observation: the weight update matrix ΔW\Delta W during fine-tuning has low intrinsic rank. If ΔW∈Rd×d\Delta W \in \mathbb{R}^{d \times d} but rank(ΔW)≈r≪d\text{rank}(\Delta W) \approx r \ll d, then we can represent it as ΔW=BA\Delta W = BA where B∈Rd×rB \in \mathbb{R}^{d \times r} and A∈Rr×dA \in \mathbb{R}^{r \times d}. This reduces the parameter count from d2d^2 to 2dr2dr — a factor of d/(2r)d/(2r) savings. For a typical d=4096d = 4096 and r=16r = 16, this is a 128x reduction.

Chapter Roadmap

Click any topic to jump in

1
Why PEFT

The memory wall, catastrophic forgetting, and intrinsic dimensionality — why full fine-tuning is impractical for large VLMs.

Three PEFT approaches

Low-rank, bottleneck, and soft prompt methods

2
LoRA

Low-rank matrix decomposition for weight updates — the most popular PEFT method, training <1% of parameters.

3
Adapter Tuning

Small bottleneck modules inserted between layers — an alternative to LoRA with different composition properties.

4
Prefix & Prompt Tuning

Learnable soft tokens prepended to input — the lightest PEFT methods, tuning only thousands of parameters.

Quantization meets LoRA
5
QLoRA

Combining 4-bit quantization with LoRA — fine-tuning 65B models on a single 48GB GPU.

From training to deployment

How to fine-tune and serve in production

6
VLM Fine-Tuning Practice

Component-wise strategies, hyperparameter selection, and common failure modes when fine-tuning VLMs.

7
Merging & Serving

Weight merging, multi-adapter serving, and TIES-Merging — deploying fine-tuned models efficiently.

Sign up to unlock this chapter

This chapter is part of PixelBank Premium. Create a free account, then upgrade to read the full lesson — concepts, walkthroughs, and exercises.