PIXELBANKv9.1.0
Menu
Back to LLM Study Plan
Week 3-4

Chapter 4: Pretraining

Understand how LLMs are trained from scratch on massive text corpora. Learn about language modeling objectives (causal vs masked), the datasets and data pipelines that fuel pretraining, the distributed training infrastructure required, and the compute-optimal strategies that determine how to allocate your training budget.

Chapter Overview

Pretraining is the most expensive and foundational phase of building an LLM. During pretraining, the model learns language structure, factual knowledge, reasoning patterns, and coding ability by predicting tokens across trillions of training examples. The cost can range from thousands to hundreds of millions of dollars.

The key insight behind modern pretraining is remarkably simple: train a large transformer to predict the next token in a sequence, and let scale do the rest. This simple objective, applied to enough data with enough parameters, produces models capable of translation, summarization, code generation, mathematical reasoning, and more---none of which were explicitly trained for.

However, the simplicity of the objective belies the complexity of the engineering. Pretraining requires: carefully curated and deduplicated datasets, distributed training across thousands of GPUs, sophisticated optimization strategies, and careful monitoring for training instabilities. This chapter covers all these aspects.

This chapter covers:

  • Language Modeling Objectives: How next-token prediction and masked prediction differ
  • Causal vs Masked LM: When to use each, and why causal dominates for generation
  • Training Data: Where it comes from, how it's cleaned, and why data quality matters more than quantity
  • Training Infrastructure: Distributed training, mixed precision, and parallelism strategies
  • Scaling Laws & Compute: Optimizing the allocation of your training budget

Chapter Roadmap

Click any topic to jump in

1
LM Objectives

Cross-entropy loss, perplexity, and teacher forcing — the training objectives that drive next-token prediction in every LLM.

Cross-Entropy LossPerplexityTeacher ForcingToken-Level vs Sequence-Level Objectives
Choosing the right objective shapes data needs

Objectives determine architecture and data

2
Causal vs Masked LM

Autoregressive vs bidirectional objectives — why causal modeling dominates for generation and how BERT's masking works.

Causal Language Modeling (CLM)Masked Language Modeling (MLM)Prefix Language ModelingWhy Causal LM Dominates
3
Training Data

Datasets, deduplication, and data mixtures — the fuel that determines what an LLM knows and how well it reasons.

Common Training DatasetsData Cleaning PipelineData DeduplicationData Mixture and Curriculum
Scaling the training pipeline

Infrastructure and theory for training at scale

4
Training Infrastructure

Data parallelism, model parallelism, and mixed precision — the distributed systems engineering behind pretraining.

Data ParallelismModel Parallelism (Tensor & Pipeline)Mixed Precision TrainingOptimizer Choices
5
Scaling Laws

Chinchilla-optimal compute allocation — predicting loss from parameters, data, and FLOPs before training begins.

The Compute BudgetLoss Prediction from ComputeCompute-Optimal AllocationBeyond Chinchilla: Inference-Aware Scaling

A transformer block is only an architecture. Before training, its billions of weights are random and it predicts nothing. To learn from raw text with no human labels, we need a task that the text itself can grade, and a single number that says how wrong the model is at every position. The choice of that task decides what the model can later do, and the choice of that number decides how smoothly it learns.

The previous chapter, Transformer Architecture, built the network. This chapter, Pretraining, covers how that network is trained on trillions of tokens, starting with the objective.

We start with cross-entropy loss, the negative log probability of the true next token. Then we turn loss into perplexity, a number that is easier to interpret. Next comes teacher forcing, the trick that lets one forward pass train every position at once, and the exposure bias it creates. We finish by contrasting token-level objectives, used in pretraining, with sequence-level ones, used later in RLHF.

Definition

A language modeling objective is the self-supervised training loss that turns raw text into a learning signal. For an autoregressive model it is the average cross-entropy between the model's predicted distribution over the vocabulary and the actual next token, taken over every position. Minimising it maximises the likelihood the model assigns to the training text, and its exponential is perplexity.

In this topic

1Cross-Entropy Loss
2Perplexity
3Teacher Forcing
4Token-Level vs Sequence-Level Objectives
1 of 4
Cross-Entropy Loss

L=−1T∑t=1Tlog⁡P(xt∣x<t;θ)\mathcal{L} = -\frac{1}{T} \sum_{t=1}^{T} \log P(x_t \mid x_{<t}; \theta)

At each position tt the model outputs a probability distribution over the vocabulary, given the earlier tokens x<tx_{<t} and weights θ\theta. The loss for that position is −log⁡P(xt∣x<t;θ)-\log P(x_t \mid x_{<t}; \theta), the negative log of the probability given to the token that actually came next. Averaging over all TT positions gives the formula. A confident correct prediction costs almost nothing, and a confident wrong one costs a lot, because −log⁡p-\log p grows without bound as pp goes to 0. Loss is measured in nats with the natural log. Its value depends on the tokenizer, so losses from different vocabularies are not directly comparable.

Mathematical Intuition

The loss L=−1T∑tlog⁡P(xt∣x<t)\mathcal{L} = -\frac{1}{T}\sum_t \log P(x_t \mid x_{<t}) is a maximum likelihood estimator — minimizing it is equivalent to minimizing DKL(Pdata∥Pθ)D_{\text{KL}}(P_{\text{data}} \| P_\theta). With a vocabulary of VV tokens and context length TT, the model must assign probability mass across VTV^T possible sequences. A loss of 2.0 means the model's average per-token surprise is 2 nats, equivalent to choosing among e2≈7.4e^2 \approx 7.4 equally likely tokens. Each 0.1 reduction in loss multiplies the probability assigned to correct tokens by e0.1≈1.1e^{0.1} \approx 1.1 — a 10% improvement in confidence.

Example:

A model assigns probability 0.8 to the correct next token. What is the per-token loss?

2 of 4
Perplexity

PPL=eL=exp⁡(−1T∑t=1Tlog⁡P(xt∣x<t))\text{PPL} = e^{\mathcal{L}} = \exp\left(-\frac{1}{T} \sum_{t=1}^{T} \log P(x_t \mid x_{<t})\right)

Perplexity is the exponential of the average cross-entropy, as in the formula. If a model's loss is L\mathcal{L} nats, its perplexity is eLe^{\mathcal{L}}. It reads as an effective branching factor: a perplexity of 10 means the model is as uncertain, on average, as if it were choosing uniformly among 10 tokens. Lower is better, and a perfect model would reach 1. Because it exponentiates the loss, small loss changes become visible ratios. GPT-3 reached a zero-shot perplexity of about 20 on Penn Treebank. Perplexity depends on the tokenizer and dataset, so compare it only between models that share both.

Mathematical Intuition

Perplexity PPL=eL\text{PPL} = e^{\mathcal{L}} converts nats to an effective vocabulary size. If PPL = 20, the model's uncertainty at each step is equivalent to a uniform distribution over 20 tokens. The key insight: perplexity is multiplicative — reducing PPL from 20 to 10 is the same relative improvement as 10 to 5. This is why we compare log-perplexity (cross-entropy) for arithmetic differences. For a vocabulary of V=50,000V = 50{,}000, random guessing gives PPL = 50,000; a well-trained model achieving PPL = 15 has reduced uncertainty by a factor of 3,333.

Example:

Model A has perplexity 15 on a test set. Model B has perplexity 10. How much better is B?

3 of 4
Teacher Forcing

During training, the input at every position is the true text, not the model's own earlier predictions. This is teacher forcing. Because every position's input is known in advance, one forward pass with a causal mask computes the loss at all TT positions in parallel, which is what makes pretraining on trillions of tokens feasible. The cost is a mismatch with generation, where the model feeds on its own samples. It never practised recovering from its own mistakes, so one bad token can push it into text unlike anything it saw. This is called exposure bias, and sampling strategies partly compensate for it.

Mathematical Intuition

Teacher forcing computes all TT positions in parallel: the loss gradient ∇θL\nabla_\theta \mathcal{L} provides a signal at every token simultaneously, giving O(T)O(T) gradient information per sequence versus O(1)O(1) for sequence-level methods. The train-test mismatch (exposure bias) grows with generation length: after kk autoregressive steps, errors compound as ϵk≤k⋅ϵ1\epsilon_k \leq k \cdot \epsilon_1 in the worst case, where ϵ1\epsilon_1 is the single-step error rate. For a 512-token generation, even a 0.1% per-token error rate compounds to ~40% chance of at least one error propagating.

Example:

Training sequence: "The cat sat on the mat." If the model predicts "dog" instead of "cat" at position 2, what does it see at position 3?

4 of 4
Token-Level vs Sequence-Level Objectives

A token-level objective scores each next-token prediction separately against the known text, so every position gives a gradient and the signal is dense and low-variance. A sequence-level objective scores a whole generated output, for example with a reward model, a human preference, or a test that runs generated code. That matches what users care about, but the model must first generate text, the score is a single number for many tokens, and the gradient must be estimated with methods such as REINFORCE or PPO. So the standard recipe is token-level pretraining followed by sequence-level fine-tuning with RLHF.

Mathematical Intuition

Token-level loss provides TT gradient signals per sequence (one per position), while sequence-level objectives provide just 1 signal for the entire sequence. The variance of policy gradient estimators scales as O(VT)O(V^T) for sequence-level rewards (exponential in length), versus O(V)O(V) for per-token cross-entropy. This is why pretraining uses token-level loss: with T=2048T = 2048 and V=50,000V = 50{,}000, the search space is 50,000204850{,}000^{2048} — far too large for sequence-level exploration. Sequence-level objectives become practical only after pretraining narrows the distribution to a manageable subspace.

Example:

Why not use sequence-level objectives during pretraining?

Theory Exercise

Problem:

A model is pretrained with cross-entropy loss and achieves perplexity 12 on a held-out test set. However, when used as a chatbot, it gives poor responses. Why might a low perplexity not translate to good conversational ability?

Hints:
  • Think about what the training data looks like vs what a chatbot should do
  • Consider the distribution of text on the internet
  • Think about the difference between predicting text and following instructions

Related Problems on PixelBank