Chapter 1: Introduction to LLMs
Understand what Large Language Models are, how they evolved from simple n-gram models to billion-parameter transformers, and why scaling changed everything. Explore key architectures, real-world applications, and the fundamental limitations that shape modern AI research.
Chapter Overview
Large Language Models (LLMs) represent a paradigm shift in artificial intelligence. Rather than hand-coding rules for language understanding, we train massive neural networks on vast corpora of text, and they learn to generate, translate, summarize, and reason about language with remarkable fluency.
The story of LLMs is one of scale. Early language models used statistical methods like n-grams and hidden Markov models. The introduction of neural language models (Bengio et al., 2003), followed by recurrent architectures (LSTMs, GRUs), and ultimately the Transformer (Vaswani et al., 2017) set the stage. But it was the discovery that simply scaling up model size, data, and compute leads to predictable capability gains---the so-called scaling laws---that ignited the modern LLM revolution.
Today, models like GPT-4, Claude, Llama, and Gemini power applications from code generation to medical diagnosis. Understanding how these systems work, what they can and cannot do, and where the field is heading is essential for any ML practitioner.
This chapter covers:
- What are LLMs? Defining the class of models and what makes them "large"
- History of Language Models: From n-grams to transformers
- Key Architectures: Encoder-only, decoder-only, and encoder-decoder designs
- Scaling Laws: How performance improves predictably with scale
- Applications: Real-world uses across industries
- Limitations: Hallucinations, reasoning gaps, and alignment challenges
Chapter Roadmap
Click any topic to jump in
What Are LLMs
Language modeling as next-token prediction — how a simple objective produces emergent intelligence at scale.
How language models evolved into distinct architectural families
History of Language Models
From n-gram counting to neural nets to transformers — each era solved the previous era's bottleneck.
Key Architectures
Encoder-only, decoder-only, encoder-decoder, and MoE — the design space of modern LLMs.
Scaling Laws
Power-law relationships between model size, data, compute, and loss — the physics of training LLMs.
Real-world impact and fundamental constraints
Applications
Code generation, conversation, extraction, and scientific discovery — where LLMs deliver real value.
Limitations
Hallucination, reasoning gaps, knowledge cutoff, and bias — the boundaries of current LLM capabilities.
Suppose you want a program that can answer questions, summarise a contract, translate a paragraph, and fix a bug in a Python function. For decades each of these was a separate system with its own hand-built features, labelled dataset, and model. Building all four meant four research projects, and none of them helped the others. A single system that does all of this from one training run would have sounded like science fiction in 2015.
Large Language Models are that system, and the surprise is how plain their training objective is. They read enormous amounts of text and learn to guess the next token. This is the first topic of the course, so it sets the vocabulary that every later chapter relies on.
We start with next-token prediction and the probability model behind it. Then we measure what makes a model large, in parameters, data, and compute, and look at emergent capabilities that appear only at scale. We finish with foundation models, the idea that one pretrained network can be adapted cheaply to thousands of downstream tasks.
Definition
A Large Language Model (LLM) is a neural network, almost always a Transformer, with billions of parameters, trained on trillions of tokens to estimate the conditional probability of the next token given all previous tokens. Text is generated by repeatedly sampling from that distribution. Its general abilities, and its adaptability to new tasks, come from this single self-supervised objective applied at very large scale.
In this topic
Language Modeling as Next-Token Prediction
Every capability of an LLM rests on one learned function: a probability distribution over the next token. The model reads the context through its layers and produces a hidden state of dimension . The output matrix , of shape where is the vocabulary size, turns into one score per token, called a logit, and softmax turns those scores into probabilities that sum to 1. Training raises the probability of the token that actually came next. The failure mode is that the model reproduces whatever its data made likely, which is not the same as what is true.
The chain rule decomposes , meaning any joint distribution over sequences can be modeled autoregressively. The softmax output maps a -dimensional hidden state to a -dimensional probability simplex. Training minimizes cross-entropy: , which is equivalent to minimizing KL divergence from the true data distribution.
Given the context "The capital of France is", what does the LLM compute?
What Makes an LLM 'Large'
Three quantities define scale, and they are linked. The parameter count runs from 117 million in GPT-1 to hundreds of billions today. The training data is measured in tokens, from hundreds of billions to many trillions. The compute is measured in floating-point operations, and for a dense Transformer a good rule is , because each token costs about operations forward and backward. Increasing only one of the three wastes money, since a huge model trained on too little data never uses its capacity. That imbalance is exactly what the Scaling Laws topic quantifies.
The parameter count typically scales as for a transformer with layers and dimension . Doubling quadruples parameters. The compute budget follows FLOPs where is training tokens. A 70B model trained on 2T tokens requires approximately FLOPs — thousands of GPU-years.
GPT-3 has 175B parameters trained on 300B tokens. Llama 2 has 70B parameters trained on 2T tokens. Which factors differ?
Emergent Capabilities
Some abilities seem to switch on suddenly as models grow. On tasks such as multi-digit arithmetic or word unscrambling, small models score near zero, and above some scale the score jumps. Wei et al. (2022) called these emergent abilities and listed in-context learning, chain-of-thought reasoning, and instruction following among them. Part of the jump is real and part is a measurement effect. Exact-match scoring gives no credit for an answer that is almost right, so steady improvement on each step can look like a sudden leap in the final score. When you read an emergence claim, check which metric produced it.
Emergence can be formalized as a phase transition in performance: below a threshold scale , task accuracy , and above it jumps sharply. Some researchers argue this is an artifact of nonlinear evaluation metrics — under log-linear metrics, performance often improves smoothly. The debate centers on whether is truly discontinuous or just appears so under certain metrics.
A 1B parameter model cannot do 3-digit addition. A 100B model can. Why?
Foundation Models
Before 2018 most NLP systems were trained from scratch for a single task. A foundation model reverses this: one large model is pretrained once on broad data and then adapted to many downstream tasks by prompting, fine-tuning, or small add-on modules. Bommasani et al. (2021) coined the term to stress both the leverage and the risk. Pretraining cost is paid once and shared by every application, but any bias or flaw in the base model is inherited by everything built on it. Parameter-efficient methods such as LoRA push the adaptation cost down further by training only a small low-rank update.
Transfer learning exploits the factorization . Fine-tuning adjusts the ratio term with far fewer examples than learning from scratch. LoRA approximates the weight update as where , , with rank , reducing trainable parameters from to .
Why is it more efficient to fine-tune a foundation model than train from scratch for each task?
Theory Exercise
Problem:
Explain why next-token prediction, a seemingly simple objective, leads to models that can perform complex tasks like writing code or solving math problems.
Hints:
- Think about what a model must understand to predict the next token accurately
- Consider the diversity of text in training data
- Think about what it means to predict the next token in a math proof
Related Problems on PixelBank
Why did it take until the late 2010s for machines to write fluent paragraphs, when researchers had been building language models since the 1940s? Each generation of models hit a specific wall. Counting-based models could not see past a few words. Early neural models generalised better but were slow and still short-sighted. Recurrent networks could in principle read a whole document, yet in practice they forgot its beginning and could not be trained in parallel on GPUs.
The previous topic, What are LLMs?, defined the goal as estimating the next token given the context. This topic follows the attempts to reach that goal, because every design choice in a modern LLM is a repair of an earlier failure.
We start with n-gram models and the sparsity problem. Next come neural language models, which introduced word embeddings, and then recurrent networks and LSTMs, whose gates eased vanishing gradients. We end with the 2017 Transformer, which replaced recurrence with attention. Its parallelism made training on web-scale data affordable.
Definition
A language model assigns a probability to a sequence of words by factoring it into next-word conditionals. Historically these conditionals were estimated in four successive ways. First came n-gram counts over the last words, then feedforward networks over learned embeddings. Recurrent networks followed, carrying a hidden state through the sequence. Transformers now use self-attention over the full context, computed in parallel.
In this topic
N-gram Models
The first practical language models simply counted. An n-gram model assumes the next word depends only on the previous words, so the full history is truncated to a short window. Each conditional is estimated as a ratio of counts: how often the window was followed by , divided by how often the window appeared at all. A bigram model () looks back exactly one word. The approach is fast and needs no training, but it has two failure modes. It cannot see beyond its window, and any word sequence absent from the corpus gets probability zero unless smoothing redistributes mass.
An -gram model with vocabulary has possible -grams. For and : entries — impossible to store. Even with aggressive pruning, the data sparsity problem means most -grams are never observed. Smoothing techniques like Kneser-Ney redistribute probability mass: .
In a bigram model trained on a corpus, P("the" | "is") = 0.15. What does this mean?
Neural Language Models
Bengio et al. (2003) attacked the sparsity problem by replacing count tables with a function. Each word gets a learned dense vector, an embedding, of perhaps 100 dimensions. The model concatenates the embeddings of the previous few words, passes them through a hidden layer, and applies softmax over the vocabulary. Because similar words end up with similar vectors, evidence about one word transfers to its neighbours, which count tables cannot do. The cost was speed. The softmax over a large vocabulary made training slow for years, and the context window was still fixed, just as in the n-gram models before it.
Bengio's model computes where are learned embeddings. The embedding matrix has only parameters instead of , and similar words share similar embeddings, enabling generalization to unseen word combinations via the smoothness of the neural network function.
Why do neural language models generalize better than n-gram models?
Recurrent Neural Networks (RNNs & LSTMs)
Recurrent networks removed the fixed window. At each step the hidden state is computed from the previous state and the current input using shared weight matrices and , a bias , and a nonlinearity . In principle summarises the entire history. In practice, gradients passed back through many steps are multiplied by similar factors each time and shrink toward zero, so early tokens stop influencing learning. LSTMs (Hochreiter and Schmidhuber, 1997) added a gated cell state that information can flow through almost unchanged. They dominated NLP from about 2013 to 2017.
The LSTM gating mechanism controls information flow: forget gate , input gate , cell update , new cell . The gradient through the cell state is , which stays near 1 when the forget gate is open, solving the vanishing gradient problem that plagues vanilla RNNs.
Why do standard RNNs struggle with long sequences?
The Transformer Revolution (2017)
Recurrence has an unavoidable bottleneck: step cannot start until step has finished, so a GPU sits mostly idle. The Transformer of Vaswani et al. (2017) dropped recurrence entirely. Self-attention lets every position look directly at every other position in one matrix multiplication, so a whole sequence is processed in parallel and any two tokens are one step apart, which eases vanishing gradients over distance. The price is memory and compute that grow with the square of sequence length. Parallel training made web-scale datasets affordable, and every model later in this course builds on this architecture.
Self-attention computes operations but all in parallel, while RNNs compute operations sequentially. For a 1000-token sequence with : transformer attention takes operations in one parallel step; an RNN takes operations but in 1000 sequential steps. The transformer's parallelism is why GPU utilization jumps from ~30% (RNNs) to ~70%+ (transformers).
Why can transformers train faster than RNNs on the same data?
Theory Exercise
Problem:
Compare the computational complexity of processing a sequence of length with an RNN versus a Transformer. What are the tradeoffs?
Hints:
- Think about sequential vs parallel computation
- Consider the self-attention matrix size
- Think about what happens as sequence length grows
A sentiment classifier, a chatbot, and a translation system all start from a Transformer, yet they need different things from it. The classifier wants to read the whole sentence at once, including words after the one it is judging. The chatbot must write text one token at a time without peeking at words it has not produced yet. The translator must fully understand one sentence before writing a different one. A single fixed wiring cannot serve all three equally well.
The previous topic, History of Language Models, ended with the 2017 Transformer. Within a year researchers had split it into distinct families, and the choice of family still decides what a model can and cannot do.
We cover encoder-only models such as BERT, which see context in both directions and excel at understanding. Next come decoder-only models such as GPT, which see only the past and excel at generation. Encoder-decoder models such as T5 join the two for sequence-to-sequence tasks. We finish with Mixture of Experts, which grows parameters without growing per-token compute.
Definition
An LLM architecture is defined by its attention masking and its pretraining objective. Encoder-only models use bidirectional attention and masked-token prediction. Decoder-only models use a causal mask and next-token prediction. Encoder-decoder models combine a bidirectional encoder with a causal decoder that attends to it through cross-attention. Mixture of Experts replaces dense feedforward layers with routed, sparsely activated experts.
In this topic
Encoder-Only (BERT-style)
An encoder-only model lets every token attend to every other token, left and right. It is pretrained with masked language modelling: about 15 percent of positions are hidden, and the model predicts each hidden token from all the others, as in the formula. Seeing both sides makes its representations strong for understanding tasks such as classification, named-entity recognition, and retrieval. The weakness is generation. The model was never trained to produce text left to right, and its objective assumes the surrounding tokens already exist. It also learns only from the masked positions, so each sequence yields fewer training targets. Examples include BERT, RoBERTa, and DeBERTa.
BERT's masked language model objective maximizes where is the set of masked positions. Because each masked token conditions on the full bidirectional context, the model captures pairwise interactions per layer. However, this objective cannot be directly used for generation because it assumes all non-masked tokens are known.
Why can't encoder-only models generate text autoregressively?
Decoder-Only (GPT-style)
A decoder-only model applies a causal mask so position can attend only to positions 1 through . It is trained to predict each next token, which by the chain rule is exactly the factorisation in the formula: the probability of a sequence is the product of next-token conditionals. Every position supplies a training target, and the same network generates text by sampling one token and appending it. With enough scale and instruction tuning, a decoder can also handle understanding tasks through prompting. That versatility is why GPT, Claude, Llama, and Mistral are all decoders. The cost is one-directional context within each forward pass.
The autoregressive factorization is mathematically exact by the chain rule — no approximation needed. The causal mask is a lower-triangular matrix: . Applied before softmax as , this zeros out attention to future tokens, ensuring the model can only use past context for prediction.
Why have decoder-only models become dominant over encoder-only for most tasks?
Encoder-Decoder (T5-style)
An encoder-decoder model splits the work. The encoder reads the entire input bidirectionally and produces one vector per input token. The decoder generates the output causally and, in each layer, uses cross-attention: its queries come from the partial output, while keys and values come from the encoder's vectors. This fits tasks where input and output are different sequences, such as translation or summarisation, because the source is fully understood before any word is written. T5 casts every task, even classification, as text-to-text. The costs are two stacks to run and a less natural fit for open-ended chat. BART and Flan-T5 are other examples.
Cross-attention in the decoder computes where queries come from the decoder and keys/values come from the encoder. This creates an information bottleneck of size — the encoder must compress the entire source into representations that the decoder's cross-attention can retrieve from. The separation allows the encoder to build bidirectional source representations.
For English-to-French translation, what does the encoder see vs the decoder?
Mixture of Experts (MoE)
A Mixture of Experts layer replaces one large feedforward block with smaller experts and a gating network . For each token, scores the experts. Only the top- experts, usually 1 or 2, are run, and their outputs are combined using the gate weights . Total parameters grow with , but each token pays only for experts, so knowledge capacity grows faster than compute. The failure mode is routing collapse, where a few experts receive most tokens and the rest go unused, so training adds a load-balancing loss. Every expert must also stay in memory. Mixtral 8x7B and Switch Transformer use this design.
The gating function routes each token to the top- experts. For experts with top-2 routing: active parameters per token of total. The load balancing loss penalizes uneven routing, where is the fraction of tokens routed to expert and is the average gating probability for expert . Without this, most tokens collapse to one or two experts.
Mixtral 8x7B has 8 experts of 7B each. How many parameters are active per token?
Theory Exercise
Problem:
You need to build: (A) a sentiment classifier, (B) a chatbot, (C) a translation system. Which architecture would you choose for each and why?
Hints:
- Think about which tasks need bidirectional context vs generation
- Consider whether each task is understanding-focused or generation-focused
- Think about whether input and output are in different formats
Training a frontier model costs tens of millions of dollars in compute, and you get one shot. Should the budget go into a bigger model or into more data? Should you train for longer or stop early? Before 2020 these choices were made by intuition, and some very expensive models turned out to be badly balanced. What labs needed was a way to predict a large model's loss from cheap experiments on small ones.
The previous topic, Key Architectures, described what shape a model takes. This topic asks how big each part should be, given a fixed budget of compute.
We begin with the Kaplan scaling laws, which showed that loss falls as a smooth power law in parameters, data, and compute. Then comes the Chinchilla result, which corrected the balance between model size and data and found about 20 tokens per parameter to be compute-optimal. We work through compute-optimal allocation using the approximation . We close with inference-time scaling, where extra compute at answer time buys accuracy without retraining.
Definition
Scaling laws are empirical power-law relationships between a language model's test loss and its parameter count , training tokens , and training compute . They let small experiments predict large-model loss. Compute-optimal training, as in Chinchilla, chooses and to minimise loss at fixed , and gives roughly 20 tokens per parameter.
In this topic
Kaplan Scaling Laws (OpenAI, 2020)
Kaplan et al. trained hundreds of models and found that test loss follows a power law in each resource when the other two are not the bottleneck. Here is parameters, is dataset tokens, and is compute. The constants , , and set the scale, and the exponents, about , , and , set the slope. On a log-log plot each law is a straight line across many orders of magnitude, so small runs can predict large ones. The small exponents mean steep diminishing returns. Kaplan's recipe favoured large models trained on relatively little data, a balance Chinchilla later corrected.
The power law with means loss decreases as the power of parameters. Taking the derivative: , which shows diminishing marginal returns — each additional parameter contributes less. Doubling parameters reduces loss by a factor of , a mere 5.1% improvement for 2x the cost.
If doubling model parameters reduces loss by 5%, how much would 10x parameters reduce it?
Chinchilla Scaling Laws (DeepMind, 2022)
Hoffmann et al. fitted loss as a function of both and and asked which split minimises loss at a fixed compute budget. The answer is that both should grow as the square root of compute, and , which works out to about 20 training tokens per parameter. By that measure GPT-3, with 175B parameters and 300B tokens, used only 1.7 tokens per parameter and was badly undertrained. Chinchilla itself had 70B parameters and 1.4T tokens. It used a similar budget to the 280B-parameter Gopher and beat it on most benchmarks.
Chinchilla minimizes subject to the compute constraint . Using Lagrange multipliers: the optimal ratio is tokens per parameter. This means GPT-3 (175B params, 300B tokens, ratio 1.7) used 12x too few tokens. The same compute spent on a 70B model with 1.4T tokens would achieve lower loss.
You have budget for 10^23 FLOPs. Chinchilla-optimal model size?
Compute-Optimal Training
For a dense Transformer, training compute is well approximated by , where is the parameter count and is the number of training tokens. Fix , and every choice of decides , so the real question is where on that curve loss is lowest. Chinchilla's fit puts the optimum near 20 tokens per parameter. A model far above that ratio is under-sized, and one far below is undertrained. There is an important caveat. Compute-optimal ignores inference. A model served to millions of users is often deliberately trained well past 20 tokens per parameter, as Llama was, because a smaller model is cheaper on every request.
Given compute budget , substituting the Chinchilla optimal gives , so and . Both scale as , confirming equal allocation. The key insight: a model that is 10x cheaper to train but trained on 10x more data (with x fewer parameters) achieves similar loss.
Budget: train either (A) 70B model on 1.4T tokens or (B) 175B model on 560B tokens. Same compute. Which is better?
Inference-Time Scaling
Compute can also be spent when answering rather than when training. The model can write a long chain of thought before answering. It can sample independent solutions and take a majority vote, which is called self-consistency. It can also search over partial reasoning paths with a verifier. None of this changes the weights, so the model's knowledge is fixed, but it gets more chances to use that knowledge correctly. Gains follow a diminishing-returns curve, similar to training-time scaling. Voting fails when mistakes are correlated, because if the model reliably makes the same error, extra samples only repeat it.
Self-consistency with samples and majority voting: if single-sample accuracy is , the majority vote accuracy is . For and : . This follows from the binomial CDF. The improvement is logarithmic in , giving diminishing returns — the same cost-performance tradeoff as training-time scaling.
A model with 5 seconds of thinking gets 60% on math. With 60 seconds of thinking (more inference compute), it gets 85%. What changed?
Theory Exercise
Problem:
A lab has a fixed compute budget and must choose between: (A) Training a 13B parameter model on 260B tokens, or (B) Training a 7B parameter model on 500B tokens. Using Chinchilla scaling laws, which is better?
Hints:
- Calculate the tokens-per-parameter ratio for each option
- The Chinchilla-optimal ratio is ~20 tokens per parameter
- Consider which is closer to the optimal ratio
A general model that can write, read, and reason about text is only valuable if it solves real problems better or cheaper than the alternatives. Some deployments have delivered: code assistants now draft a large share of new code at many companies. Others have failed publicly, with invented legal citations and chatbots promising refunds that did not exist. The difference is rarely raw model quality. It is whether the application checks the model's output, and whether errors are cheap to catch.
The previous topic, Scaling Laws, explained why larger models trained on more data keep getting better. This topic asks where that capability pays off in practice.
We begin with code generation, where output can be tested automatically, which makes it the strongest use case. Then we cover conversational assistants and the instruction tuning that turns a text predictor into a helpful agent. Next come information extraction and summarisation at scale, where volume is the win and recall is the risk. We end with scientific research, where retrieval over large literatures helps but every claim must be verified against sources.
Definition
An LLM application wraps a pretrained model in a system of prompts, retrieved context, tools, and output checks to perform a specific task. Its reliability depends on the model and on how cheaply errors can be detected. Examples include test suites for code, schemas for extraction, human review for high-stakes answers, and source citations for research claims.
In this topic
Code Generation & Software Engineering
Code is the application where LLMs have been most useful, because correctness can be checked by machine. Models trained on large code corpora, such as Codex and its successors behind GitHub Copilot and Claude Code, complete functions, write tests, explain unfamiliar code, and translate between languages. The Codex paper introduced pass@k, the probability that at least one of sampled programs passes the unit tests. It exposed a useful property: sampling several attempts and keeping the one that passes beats trusting a single attempt. The failure mode is plausible code that compiles but is subtly wrong, or uses an API that does not exist, whenever tests are weak.
Code generation can be viewed as constrained text generation where the output must satisfy formal grammar rules: only if the token produces a valid partial parse. In practice, LLMs approximate this soft constraint by assigning near-zero probability to syntax errors after sufficient training on syntactically correct code. The tree-sitter grammar can be used for constrained decoding, guaranteeing valid output.
How does an LLM 'understand' code differently from natural language?
Conversational AI & Assistants
A pretrained model continues text. It does not answer questions. Assistants such as ChatGPT and Claude are made by two further training stages. Supervised fine-tuning (SFT) shows the model thousands of written examples of good assistant replies. Reinforcement learning from human feedback (RLHF) then trains a reward model on human preferences between pairs of responses and optimises the model toward higher reward. This covers instruction following, multi-turn context, and refusal of harmful requests. A context window limits how much conversation the model can see at once, and a longer chat eventually pushes early turns out of view.
RLHF optimizes where is the learned reward model. The KL penalty prevents the policy from drifting too far from the reference (pretrained) model. The reward model is trained on human preference comparisons: given responses , maximize .
Why is a raw pretrained LLM poor at conversation compared to ChatGPT?
Information Extraction & Summarization
Many organisations sit on huge piles of unstructured text: contracts, clinical notes, filings, and support tickets. An LLM can read each document and fill a fixed schema of fields, such as parties, dates, amounts, and clause types, or produce a summary. Asking for structured output, such as JSON that matches a schema, makes results machine-checkable and easy to aggregate. The value is speed at volume. The main risk is recall. A model that silently misses a clause looks exactly like one that found nothing, so production pipelines measure recall on a labelled sample and route low-confidence documents to human reviewers.
Extraction can be formulated as structured prediction: given document , produce structured output where each is a key-value pair. The model learns where the schema specifies expected fields. The attention mechanism naturally performs soft information retrieval: attention weights find relevant spans in the document for each output field.
A law firm needs to review 10,000 contracts for specific clauses. How can LLMs help?
Scientific Research & Discovery
Research literatures have outgrown what any person can read. LLMs help by summarising papers, extracting findings into tables, suggesting hypotheses, and drafting analysis code. Domain models such as BioGPT and Med-PaLM are trained or tuned on specialist corpora. The dominant pattern is retrieval-augmented generation (RAG). A search step first retrieves the relevant papers, and the model then answers using only those passages, with citations. That keeps claims tied to checkable sources. The failure mode is serious, because a fabricated citation or a misread effect size can look authoritative. Every claim must be traced back to a primary source before it is used.
RAG (Retrieval-Augmented Generation) extends the LLM's effective knowledge by conditioning on retrieved documents: where uses embedding similarity for retrieval. This sidesteps the knowledge cutoff problem by grounding generation in up-to-date sources, though it introduces a new failure mode: retrieval errors propagating into generated text.
A researcher needs to understand how a gene relates to a disease across 5,000 papers. How can LLMs help?
Theory Exercise
Problem:
Design an LLM-powered system for a customer support team that handles 1000 tickets/day. What components would you need? What could go wrong?
Hints:
- Think about the pipeline: classification, routing, response generation
- Consider what happens when the LLM is wrong
- Think about customer trust and escalation
A lawyer files a brief whose case citations were supplied by a chatbot, and the judge discovers that six of the cases do not exist. A model answers a maths question fluently and gets the arithmetic wrong. An assistant describes last year's news as current. Each failure came from a system that sounded equally confident when it was right. Deploying LLMs safely starts with knowing where and why they break.
The previous topic, Applications of LLMs, showed where these models create value. This one maps the boundaries, and each limitation follows directly from next-token prediction on a fixed corpus.
We start with hallucination, fluent text that is not grounded in fact. Next we look at reasoning limits, where pattern completion stands in for step-by-step computation. Then comes the knowledge cutoff, which freezes the model's picture of the world at training time. We end with bias and safety: the model inherits skewed patterns from its data and can be steered toward harmful output. Each section pairs the failure with the standard mitigation.
Definition
LLM limitations are systematic failure modes that arise from training a fixed network to predict likely text. They include hallucination, meaning confident statements not supported by fact or source. They also include unreliable multi-step reasoning, knowledge frozen at the training cutoff, and biased or unsafe outputs inherited from the training data. Mitigations include retrieval, tool use, and alignment training.
In this topic
Hallucination
A hallucination is fluent output that is false or unsupported, such as an invented citation, a wrong date, or a nonexistent API. It follows from the objective. Training rewards text that is likely given the context, not text that is verified, and the model has no built-in signal separating recalled facts from plausible guesses. Rare facts are the most exposed, because the model saw them only a few times and similar-looking alternatives compete. Reported rates vary widely by task and model, from a few percent on grounded summarisation to far higher on obscure questions. Mitigations include retrieval with citations, abstention training, and checking claims against sources.
Hallucination arises because the training objective rewards fluency, not factual accuracy. The model learns from co-occurrence statistics, not from a verified knowledge base. When the model encounters a context where multiple continuations are plausible, it selects the most likely token sequence — which may be factually wrong but statistically common. Calibration analysis shows models are often overconfident: .
An LLM confidently states: "The Eiffel Tower was built in 1892 and is 324 meters tall." What is wrong?
Reasoning Limitations
An LLM produces each token in one forward pass through a fixed number of layers. For problems that need many dependent steps, such as long multiplication, planning under constraints, or following a chain of logic, the answer cannot be computed inside a single token. Either the steps are written out, or the model falls back on patterns that resemble the answer. Performance also drops on problems phrased unlike the training data, even when the logic is identical. Chain-of-thought prompting helps by moving intermediate results into the text, where later tokens can use them. Tool calls, such as a calculator or code interpreter, remove the problem for computation.
LLMs approximate reasoning through pattern matching on token sequences. For arithmetic , the model must learn pairwise mappings from rather than implementing a multiplication algorithm. The number of distinct products for -digit numbers grows as , making memorization infeasible beyond small numbers. Chain-of-thought prompting helps by decomposing the computation into intermediate steps, each of which is easier to pattern-match.
Why might an LLM solve "What is 23 * 47?" correctly but fail on "What is 2347 * 4723?"?
Knowledge Cutoff & Staleness
A model's weights are a snapshot of its training data, which ends at a cutoff date. Anything after that, such as new events, papers, prices, or software versions, is invisible to it. Worse, the model often does not know where its knowledge ends. Because it learned facts as patterns rather than dated records, it may present stale information as current. The standard fix is retrieval-augmented generation (RAG). A search step fetches up-to-date documents and places them in the prompt, and the model answers from them. RAG has its own failure mode, because a poor retrieval result can lead the model to a confidently wrong answer.
An LLM's parameters encode a static snapshot of the training distribution . When the world distribution shifts to , the model's predictions degrade. RAG addresses this by conditioning on retrieved context: where contains up-to-date documents. The fundamental tradeoff: parametric knowledge (fast but static) versus retrieved knowledge (slower but current).
A model trained on data up to 2023 is asked 'Who won the 2024 Nobel Prize in Physics?' What happens?
Bias & Safety
Training text reflects the world that wrote it, including its stereotypes, imbalances, and harmful content. A model trained to reproduce likely text learns those skews as probabilities. For example, occupations become associated with genders or nationalities with traits, and the model can repeat them in hiring, lending, or medical contexts. Safety adds adversarial risks. Prompt injection hides instructions inside retrieved content, and jailbreaks coax a model past its refusals. Mitigations act at several levels: curating data, alignment training such as RLHF and Constitutional AI, output filtering, and audits that measure disparities across demographic groups.
Training data bias manifests as skewed conditional distributions: reflects historical text, not reality. Debiasing techniques modify either the data (), the model (), or the output (). Constitutional AI uses a self-critique loop: generate response critique for bias revise accept, effectively training to satisfy a set of principles without human labels for every case.
An LLM consistently associates "nurse" with "she" and "engineer" with "he". Why?
Theory Exercise
Problem:
You are deploying an LLM for medical question answering. List the top 5 risks and propose mitigations for each.
Hints:
- Think about hallucination in a medical context
- Consider liability and patient safety
- Think about edge cases the model might not handle