PIXELBANKv9.1.0
Menu
Back to LLM Study Plan
Week 1-2

Chapter 1: Introduction to LLMs

Understand what Large Language Models are, how they evolved from simple n-gram models to billion-parameter transformers, and why scaling changed everything. Explore key architectures, real-world applications, and the fundamental limitations that shape modern AI research.

Chapter Overview

Large Language Models (LLMs) represent a paradigm shift in artificial intelligence. Rather than hand-coding rules for language understanding, we train massive neural networks on vast corpora of text, and they learn to generate, translate, summarize, and reason about language with remarkable fluency.

The story of LLMs is one of scale. Early language models used statistical methods like n-grams and hidden Markov models. The introduction of neural language models (Bengio et al., 2003), followed by recurrent architectures (LSTMs, GRUs), and ultimately the Transformer (Vaswani et al., 2017) set the stage. But it was the discovery that simply scaling up model size, data, and compute leads to predictable capability gains---the so-called scaling laws---that ignited the modern LLM revolution.

Today, models like GPT-4, Claude, Llama, and Gemini power applications from code generation to medical diagnosis. Understanding how these systems work, what they can and cannot do, and where the field is heading is essential for any ML practitioner.

This chapter covers:

  • What are LLMs? Defining the class of models and what makes them "large"
  • History of Language Models: From n-grams to transformers
  • Key Architectures: Encoder-only, decoder-only, and encoder-decoder designs
  • Scaling Laws: How performance improves predictably with scale
  • Applications: Real-world uses across industries
  • Limitations: Hallucinations, reasoning gaps, and alignment challenges

Chapter Roadmap

Click any topic to jump in

1
What Are LLMs

Language modeling as next-token prediction — how a simple objective produces emergent intelligence at scale.

Language Modeling as Next-Token PredictionWhat Makes an LLM 'Large'Emergent CapabilitiesFoundation Models
Evolution and divergence

How language models evolved into distinct architectural families

2
History of Language Models

From n-gram counting to neural nets to transformers — each era solved the previous era's bottleneck.

N-gram ModelsNeural Language ModelsRecurrent Neural Networks (RNNs & LSTMs)The Transformer Revolution (2017)
3
Key Architectures

Encoder-only, decoder-only, encoder-decoder, and MoE — the design space of modern LLMs.

Encoder-Only (BERT-style)Decoder-Only (GPT-style)Encoder-Decoder (T5-style)Mixture of Experts (MoE)
Quantifying progress
4
Scaling Laws

Power-law relationships between model size, data, compute, and loss — the physics of training LLMs.

Kaplan Scaling Laws (OpenAI, 2020)Chinchilla Scaling Laws (DeepMind, 2022)Compute-Optimal TrainingInference-Time Scaling
Where LLMs work and where they fail

Real-world impact and fundamental constraints

5
Applications

Code generation, conversation, extraction, and scientific discovery — where LLMs deliver real value.

Code Generation & Software EngineeringConversational AI & AssistantsInformation Extraction & SummarizationScientific Research & Discovery
6
Limitations

Hallucination, reasoning gaps, knowledge cutoff, and bias — the boundaries of current LLM capabilities.

HallucinationReasoning LimitationsKnowledge Cutoff & StalenessBias & Safety

Suppose you want a program that can answer questions, summarise a contract, translate a paragraph, and fix a bug in a Python function. For decades each of these was a separate system with its own hand-built features, labelled dataset, and model. Building all four meant four research projects, and none of them helped the others. A single system that does all of this from one training run would have sounded like science fiction in 2015.

Large Language Models are that system, and the surprise is how plain their training objective is. They read enormous amounts of text and learn to guess the next token. This is the first topic of the course, so it sets the vocabulary that every later chapter relies on.

We start with next-token prediction and the probability model behind it. Then we measure what makes a model large, in parameters, data, and compute, and look at emergent capabilities that appear only at scale. We finish with foundation models, the idea that one pretrained network can be adapted cheaply to thousands of downstream tasks.

Definition

A Large Language Model (LLM) is a neural network, almost always a Transformer, with billions of parameters, trained on trillions of tokens to estimate the conditional probability of the next token given all previous tokens. Text is generated by repeatedly sampling from that distribution. Its general abilities, and its adaptability to new tasks, come from this single self-supervised objective applied at very large scale.

In this topic

1Language Modeling as Next-Token Prediction
2What Makes an LLM 'Large'
3Emergent Capabilities
4Foundation Models
1 of 4
Language Modeling as Next-Token Prediction

P(xt∣x1,x2,…,xt−1)=softmax(Wht)P(x_t \mid x_1, x_2, \ldots, x_{t-1}) = \text{softmax}(W h_t)

Every capability of an LLM rests on one learned function: a probability distribution over the next token. The model reads the context x1,…,xt−1x_1, \ldots, x_{t-1} through its layers and produces a hidden state hth_t of dimension dd. The output matrix WW, of shape V×dV \times d where VV is the vocabulary size, turns hth_t into one score per token, called a logit, and softmax turns those scores into probabilities that sum to 1. Training raises the probability of the token that actually came next. The failure mode is that the model reproduces whatever its data made likely, which is not the same as what is true.

Mathematical Intuition

The chain rule decomposes P(x1,…,xT)=∏tP(xt∣x<t)P(x_1, \ldots, x_T) = \prod_t P(x_t \mid x_{<t}), meaning any joint distribution over sequences can be modeled autoregressively. The softmax output P(xt∣x<t)=softmax(Wht)P(x_t \mid x_{<t}) = \text{softmax}(W h_t) maps a dd-dimensional hidden state to a VV-dimensional probability simplex. Training minimizes cross-entropy: L=−∑tlog⁡P(xt∣x<t)\mathcal{L} = -\sum_t \log P(x_t \mid x_{<t}), which is equivalent to minimizing KL divergence from the true data distribution.

Example:

Given the context "The capital of France is", what does the LLM compute?

2 of 4
What Makes an LLM 'Large'

Three quantities define scale, and they are linked. The parameter count NN runs from 117 million in GPT-1 to hundreds of billions today. The training data DD is measured in tokens, from hundreds of billions to many trillions. The compute CC is measured in floating-point operations, and for a dense Transformer a good rule is C≈6NDC \approx 6ND, because each token costs about 2N2N operations forward and 4N4N backward. Increasing only one of the three wastes money, since a huge model trained on too little data never uses its capacity. That imbalance is exactly what the Scaling Laws topic quantifies.

Mathematical Intuition

The parameter count NN typically scales as N≈12Ld2N \approx 12 L d^2 for a transformer with LL layers and dimension dd. Doubling dd quadruples parameters. The compute budget follows C≈6NDC \approx 6ND FLOPs where DD is training tokens. A 70B model trained on 2T tokens requires approximately 6×70×109×2×1012=8.4×10236 \times 70 \times 10^9 \times 2 \times 10^{12} = 8.4 \times 10^{23} FLOPs — thousands of GPU-years.

Example:

GPT-3 has 175B parameters trained on 300B tokens. Llama 2 has 70B parameters trained on 2T tokens. Which factors differ?

3 of 4
Emergent Capabilities

Some abilities seem to switch on suddenly as models grow. On tasks such as multi-digit arithmetic or word unscrambling, small models score near zero, and above some scale the score jumps. Wei et al. (2022) called these emergent abilities and listed in-context learning, chain-of-thought reasoning, and instruction following among them. Part of the jump is real and part is a measurement effect. Exact-match scoring gives no credit for an answer that is almost right, so steady improvement on each step can look like a sudden leap in the final score. When you read an emergence claim, check which metric produced it.

Mathematical Intuition

Emergence can be formalized as a phase transition in performance: below a threshold scale N∗N^*, task accuracy A(N)≈randomA(N) \approx \text{random}, and above it A(N)A(N) jumps sharply. Some researchers argue this is an artifact of nonlinear evaluation metrics — under log-linear metrics, performance often improves smoothly. The debate centers on whether A(N)A(N) is truly discontinuous or just appears so under certain metrics.

Example:

A 1B parameter model cannot do 3-digit addition. A 100B model can. Why?

4 of 4
Foundation Models

Before 2018 most NLP systems were trained from scratch for a single task. A foundation model reverses this: one large model is pretrained once on broad data and then adapted to many downstream tasks by prompting, fine-tuning, or small add-on modules. Bommasani et al. (2021) coined the term to stress both the leverage and the risk. Pretraining cost is paid once and shared by every application, but any bias or flaw in the base model is inherited by everything built on it. Parameter-efficient methods such as LoRA push the adaptation cost down further by training only a small low-rank update.

Mathematical Intuition

Transfer learning exploits the factorization Ptask(y∣x)=Ppretrained(y∣x)⋅Ptask(y∣x)Ppretrained(y∣x)P_{\text{task}}(y \mid x) = P_{\text{pretrained}}(y \mid x) \cdot \frac{P_{\text{task}}(y \mid x)}{P_{\text{pretrained}}(y \mid x)}. Fine-tuning adjusts the ratio term with far fewer examples than learning PtaskP_{\text{task}} from scratch. LoRA approximates the weight update as ΔW=BA\Delta W = BA where B∈Rd×rB \in \mathbb{R}^{d \times r}, A∈Rr×dA \in \mathbb{R}^{r \times d}, with rank r≪dr \ll d, reducing trainable parameters from d2d^2 to 2dr2dr.

Example:

Why is it more efficient to fine-tune a foundation model than train from scratch for each task?

Theory Exercise

Problem:

Explain why next-token prediction, a seemingly simple objective, leads to models that can perform complex tasks like writing code or solving math problems.

Hints:
  • Think about what a model must understand to predict the next token accurately
  • Consider the diversity of text in training data
  • Think about what it means to predict the next token in a math proof