Chapter 7: Prompt Engineering
Master the art and science of communicating effectively with LLMs. Learn zero-shot and few-shot prompting techniques, Chain-of-Thought reasoning that unlocks step-by-step problem solving, system prompts for controlling model behavior, and structured output techniques for reliable integration with downstream systems.
Chapter Overview
Prompt engineering is the practice of designing inputs that elicit desired outputs from LLMs. It is the primary interface between human intent and model behavior---no code changes, no fine-tuning, just carefully crafted text. Despite its apparent simplicity, effective prompting can dramatically improve model performance on complex tasks.
The core insight of prompt engineering is that LLMs are sensitive to how questions are framed. The same question asked differently can produce answers that vary from incorrect to expert-level. This sensitivity arises because the model's behavior is shaped by patterns in its training data: it responds differently to a casual question than to a formal academic query, differently to a vague instruction than to a precise one.
Modern prompt engineering has evolved from simple instruction-writing to systematic techniques like Chain-of-Thought (CoT) reasoning, which can improve math accuracy by 50%+, and structured output prompting, which enables reliable integration with software systems. Understanding these techniques is essential for any LLM practitioner.
This chapter covers:
- Zero-Shot Prompting: Getting results without any examples
- Few-Shot Prompting: Teaching by example with in-context demonstrations
- Chain-of-Thought: Unlocking step-by-step reasoning for complex problems
- System Prompts: Configuring model behavior, persona, and constraints
- Structured Output: Generating reliable JSON, XML, and formatted data
Chapter Roadmap
Click any topic to jump in
Zero-Shot
Clear instructions, role assignment, and task decomposition — getting results without any examples.
Few-Shot
In-context learning with demonstrations — how example selection and formatting drive model behavior.
Chain-of-Thought
Step-by-step reasoning that unlocks math and logic — zero-shot CoT, few-shot CoT, and self-consistency.
Control and reliability for deployed systems
System Prompts
Persona engineering, guardrails, and behavioral constraints — configuring model behavior at the system level.
Structured Output
JSON mode, schema specification, and constrained decoding — reliable machine-readable output from LLMs.
A trained, aligned model still does nothing useful until someone asks it for something, and the way you ask changes what you get. The same model can return a vague essay or a precise, correctly formatted answer depending only on the words in the prompt. No weights change between those two outcomes. So the first skill in using an LLM is writing a request that leaves the model as little room to guess as possible.
The previous chapter, RLHF & Alignment, trained models to follow instructions and prefer helpful answers. This chapter, Prompt Engineering, is about using that ability well, and this first topic covers the simplest case: a bare instruction with no examples.
We start with clear instructions, the habit of stating the task, audience, length, and format explicitly. Then we cover role assignment, where a persona shifts the vocabulary and depth of the answer. Next comes task decomposition, which spells out the steps a complex request needs. We finish with negative instructions, which rule out the failure modes you have already seen, and with why those work best when paired with a positive alternative.
Definition
Zero-shot prompting asks a model to perform a task from a natural-language instruction alone, with no worked input-output examples in the prompt. The model must infer the task, the expected format, and the level of detail from the instruction and its pretrained and instruction-tuned knowledge. Output quality therefore depends directly on how specific and unambiguous the instruction is.
In this topic
Clear Instructions
A model reads a prompt as the start of a document and continues it in the most likely way. A vague request such as "summarize this" is compatible with thousands of continuations, so the model picks a generic middle ground. Every detail you add removes candidates: the audience, the length, the format, the focus, and what counts as done. Good instructions name the task, the input, the constraints, and the output shape, in that order. Put long input text in clearly delimited blocks so it is not confused with the instruction. The failure mode is under-specification, and the fix is to write the prompt a careful new colleague would need.
Prompt specificity reduces output entropy: a vague prompt gives for a specific prompt , where is the conditional entropy of the model's output distribution. Each constraint in the prompt eliminates a fraction of the output space: with independent constraints each reducing options by half, the remaining space is of the original. In practice, constraints are correlated, so the reduction follows where is the average constraint strength.
Compare the outputs for: (A) "Tell me about dogs" vs (B) "List 5 key facts about domestic dog care that a first-time owner should know, in bullet points."
Role Assignment
A role, such as "You are an emergency physician", tells the model which part of its training distribution to imitate. Text written by specialists uses different vocabulary, assumes different background knowledge, and covers different risks than text written for a general audience. The role shifts all of those at once, which is why one sentence can change the depth of an answer. Roles work best when they come with the audience and the goal, not alone. They do not add knowledge the model lacks, and research on factual benchmarks finds persona gains small and inconsistent. So use roles to set register and depth, not to raise accuracy.
Role assignment shifts the model's conditional distribution from to . By Bayes' rule, this is proportional to — the model up-weights responses that are consistent with the assigned role. The effectiveness depends on how well the role is represented in pretraining data: 'expert cardiologist' activates patterns from medical texts (high-quality training signal), while 'alien from planet Zorg' has no grounding and mainly affects style. The expected quality improvement is proportional to — larger distributional shifts indicate the role is having more effect.
Ask "What causes chest pain?" with (A) no role and (B) "You are an emergency medicine physician."
Task Decomposition in Prompts
A single instruction such as "analyse this data" hides several sub-tasks, and the model may do only the first one or blend them together. Decomposition lists the steps explicitly, in the order they should run: identify trends, then anomalies, then causes, then recommendations. Numbered steps act as a checklist, so each part of the output maps to one requirement, and a missing section is easy to spot. It also sets the order of reasoning, so later steps can use earlier results. The cost is prompt length and some rigidity. For tasks with dependent stages, chaining separate calls, each with one step, gives even more control.
Decomposing a task into sequential steps reduces the effective complexity from to where is the complexity per step. This is because the model processes one step at a time, each with bounded context. The error probability also changes: for independent steps with per-step error , the overall success probability is for small . With steps and per step, overall success is ~77%. Without decomposition, the single-step error for the full task might be 50%+ because the model must implicitly manage all subtasks simultaneously.
You need a model to review a pull request. Write a zero-shot prompt with task decomposition.
Negative Instructions
Negative instructions tell the model what to avoid: no disclaimers, no jargon, no invented facts. They are the natural fix once you have seen a specific failure. They work best when paired with a positive alternative, such as "say I don't know if the context does not contain the answer". A bare prohibition leaves the model without a replacement behaviour, and naming an unwanted phrase can even prime it. Prohibitions are also soft: a long conversation or a conflicting user request can override them. For hard rules, such as never outputting personal data, add a check in code, not only a prompt instruction.
Negative instructions ('do not X') constrain the output distribution by zeroing out probability mass on undesired outputs: . The effectiveness depends on how much probability mass the model originally placed on forbidden outputs. If , then after the constraint, the remaining outputs are renormalized by . A common failure mode: the constraint is too vague ("don't be verbose") so is not well-defined, and the model cannot reliably exclude those outputs.
A model keeps adding 'I hope this helps!' to every response. How do you fix this?
Theory Exercise
Problem:
You need an LLM to extract key information from job postings. Design a zero-shot prompt that extracts: job title, company, location, salary range, required skills, and experience level. The output should be structured.
Hints:
- Think about the output format you need
- Consider edge cases (salary not listed, remote work, etc.)
- Be explicit about what to do when information is missing
Some tasks are hard to describe but easy to show. A labelling scheme with subtle boundaries, an unusual output format, or a house writing style can take paragraphs to explain and still be misunderstood. A zero-shot instruction leaves the model to guess what "concise" or "formal" or "the right label" means. What we want is a way to demonstrate the task directly, without training the model, and to have it copy the pattern on new inputs.
The previous topic, Zero-Shot Prompting, relied on instructions alone. This topic adds worked examples to the prompt, a technique GPT-3 showed at scale and called in-context learning.
We start with in-context learning itself, how a frozen model picks up a task from demonstrations, and what research says it actually learns from them. Then we cover shot count and quality, where gains grow roughly with the logarithm of the number of examples. Next comes example selection, including retrieving the examples most similar to each input. We finish with formatting consistency, because the model copies the format of the examples as faithfully as their content.
Definition
Few-shot prompting places demonstrations, each an input paired with its correct output, in the prompt before a new input. The model infers the task from the demonstrations and continues the pattern, with no gradient updates. This behaviour is called in-context learning, and is typically between 1 and a few dozen, limited by the context window and cost.
In this topic
In-Context Learning (ICL)
In-context learning is a model performing a new task from demonstrations in its prompt, with its weights unchanged. GPT-3 showed that the ability grows with model size. Nothing is trained: the demonstrations only condition the next-token distribution. Min et al. (2022) found something surprising about what the model uses. Replacing the demonstration labels with random ones barely hurt accuracy on many classification tasks. The demonstrations mainly teach the label space, the input distribution, and the format, not the input-to-label mapping. So ICL often recognises a task the model already knows rather than learning a new one, and it struggles when the task conflicts with pretraining habits.
In-context learning is a form of implicit Bayesian inference: given examples in the prompt, the model computes — it infers a latent function from the examples and applies it to the new input. The transformer's attention mechanism implements this by attending from the query to the example outputs, computing a weighted combination. The effective 'training' from examples occurs entirely in the forward pass — no gradient updates, just context-dependent activation patterns.
Provide a 3-shot prompt for sentiment classification.
Shot Count and Quality
The formula says performance rises roughly with the logarithm of the number of shots , where is the zero-shot level and sets how fast it improves. A log curve means each doubling adds the same gain, so going from 1 to 4 shots helps as much as going from 4 to 16. The first example gives the largest jump, because it fixes the format. After a few dozen examples, gains are small and each costs context. Quality matters as much as count. Zhao et al. (2021) showed that accuracy can swing from near chance to near state of the art with the choice and order of the same examples.
Performance scales logarithmically with shot count: for examples. This means the marginal value of each additional example decreases rapidly: going from 1 to 3 shots often improves accuracy by 10-15%, but going from 5 to 15 shots may only add 2-3%. The practical limit is the context window: examples each consuming tokens leaves tokens for the actual query and response, where is the context length. The optimal balances example coverage against context space: where penalizes context consumption.
For a classification task with 10 categories, how many few-shot examples should you provide?
Example Selection Strategy
Which examples you choose matters more than how many. Diverse examples cover the sub-types and edge cases the model will face. Similar examples, retrieved for each input by embedding similarity, give the model the closest precedent. Liu et al. (2021) found retrieval-based selection beat random examples on every task they tried, with large gains on table-to-text generation and open-domain question answering. Contrastive pairs, similar inputs with different correct outputs, show where a boundary lies. Order matters too: models are biased towards the label of the last example. A common failure is examples that are all easy, which leaves the model unprepared for hard inputs.
Choosing the most informative examples is a combinatorial optimization: from a pool of candidates, select to maximize task performance. Random selection gives a baseline; similarity-based selection (choosing examples closest to the query in embedding space: ) typically improves by 5-10%. Diverse selection (maximizing coverage of the input space) can further improve by reducing redundancy: , balancing relevance and diversity.
For a code translation task (Python to JavaScript), choose between: (A) 5 examples all translating simple functions, or (B) 5 examples covering: simple function, class, async function, error handling, and list comprehension.
Formatting Consistency
A model copies the format of its demonstrations as faithfully as their content. If one example uses Input and Output, another Question and Answer, and a third inline text, the model cannot tell which format to continue, and its output may mix them. Use one template for every example: the same labels, the same delimiters such as a blank line or three hashes, and the same level of detail. End the prompt with the template's input label filled in and its output label empty, so the continuation is forced into place. Consistent formatting also makes the output easy to parse, because every answer follows the same label.
Format consistency establishes a pattern that the model extrapolates. If examples use format , the model assigns higher probability to outputs matching : . The attention mechanism learns to copy the formatting pattern: attention weights from the output position to delimiter tokens in examples become large, effectively 'locking in' the structural template. Inconsistent formatting reduces this pattern-matching signal by splitting attention across competing format hypotheses, reducing the model's confidence in any single output structure.
What is wrong with these few-shot examples? Example 1: Input: hello -> Translation: hola Example 2: Translate 'goodbye' to Spanish: adios Example 3: Input: thank you\nOutput: gracias
Theory Exercise
Problem:
You have a named entity recognition (NER) task with 6 entity types: PERSON, ORG, LOCATION, DATE, MONEY, PRODUCT. Design a few-shot prompt that reliably extracts all entity types from arbitrary text.
Hints:
- Think about how to represent entity annotations in text
- Consider examples that cover all 6 entity types
- Think about edge cases like overlapping entities or entities with similar formats
Related Problems on PixelBank
Ask a model a multi-step word problem and ask for only the answer, and it often gets it wrong, even when it could solve every step on its own. The reason is computational. Each token is produced by a fixed number of layers, so a single answer token gets a fixed amount of computation. A problem that needs five dependent steps has to be solved inside one forward pass, with no place to store the intermediate results.
The previous topic, Few-Shot Prompting, showed the model what answers look like. This topic shows it what reasoning looks like, by having it write the intermediate steps as text before the answer.
We start with why chain-of-thought works: every generated step becomes input to the next, giving the model both more computation and a written working memory. Then we cover zero-shot chain-of-thought, the single phrase "Let's think step by step". Next comes few-shot chain-of-thought, where the examples show the reasoning as well as the answer. We finish with self-consistency, which samples several reasoning paths and takes a majority vote over their answers.
Definition
Chain-of-thought (CoT) prompting makes a model generate intermediate reasoning steps in natural language before its final answer. It is elicited either by worked examples that include their reasoning or by an instruction such as "Let's think step by step". The written steps give the model extra computation and a readable trace, which improves accuracy on arithmetic, logic, and multi-step tasks.
In this topic
Why CoT Works
The formula contrasts two routes from the input to the answer. Standard prompting must compute the answer in one forward pass, so every intermediate result has to be held implicitly inside the layers. With CoT, each step is written out as tokens, and each new token is computed with all earlier steps in its context. Writing a step stores it, and every extra token adds another full forward pass of computation. The steps also make errors visible and checkable. Wei et al. (2022) found the gains appear mainly in large models. Smaller models produce fluent but wrong reasoning, which can make accuracy worse.
Chain-of-thought decomposes the conditional distribution into . The key insight: is much higher than directly because the intermediate tokens carry computation forward. A transformer with layers and dimensions has computation per token — by generating reasoning tokens, the effective computation increases to , allowing the model to solve problems that exceed a single forward pass's capacity.
"A store sells apples for 2 dollars each. If I buy 7 apples and pay with a 20 dollar bill, how much change do I get?" Compare direct answer vs CoT.
Zero-Shot CoT
Zero-shot CoT, from Kojima et al. (2022), appends one sentence, "Let's think step by step", to the question. No examples are needed. The phrase shifts the continuation from a bare answer to a written derivation, and the model then extracts the answer from its own reasoning. On the large InstructGPT model, the paper reports accuracy rising from 17.7 to 78.7 percent on MultiArith and from 10.4 to 40.7 percent on GSM8K. It is cheap and general, but its reasoning style is uncontrolled. Recent chat and reasoning models often think step by step by default, so the gain on them is smaller.
Appending 'Let's think step by step' shifts the model's output distribution from to . This activates patterns from the pretraining data where step-by-step reasoning preceded correct answers (textbooks, tutorials, worked examples). The prompt acts as a retrieval cue: is high because this phrase co-occurred with structured reasoning in millions of training documents. Zero-shot CoT is less reliable than few-shot CoT because the model must infer both the reasoning format and the answer — two sources of uncertainty versus one.
Apply zero-shot CoT to: "If a train travels at 60 mph and needs to cover 210 miles, how long will the trip take?"
Few-Shot CoT
Few-shot CoT puts worked examples in the prompt, and each example shows the question, the reasoning steps, and the answer. This is the original method of Wei et al. (2022), who used eight hand-written exemplars for math problems. The examples teach how to reason as well as what to answer: which quantities to name, how much detail each step needs, and how to state the final answer so it can be parsed. It usually beats zero-shot CoT when the examples match the task. The costs are the effort of writing good exemplars and the extra tokens. Errors in an exemplar's reasoning get copied.
Few-shot CoT provides explicit demonstrations of the reasoning process: where is the reasoning chain. The model learns to map questions to reasoning patterns via in-context learning: . The quality of reasoning chains matters more than their correctness: even reasoning chains with wrong final answers improve performance if the intermediate steps are logically structured. This suggests CoT teaches the model a reasoning format (the 'how') more than factual knowledge (the 'what'). The improvement is most dramatic on multi-step problems where the number of reasoning steps 3.
Design a few-shot CoT prompt for percentage problems.
Self-Consistency
Self-consistency, from Wang et al. (2023), replaces the single greedy reasoning path with sampled ones. Each is generated with temperature above zero, and the final answer is the one that appears most often, as the formula's arg max over vote counts shows. The idea is that there are many ways to reach the right answer but errors scatter across different wrong ones. The paper reports gains of 17.9 points on GSM8K and 11.0 on SVAMP. The cost is times the generation. It needs answers that can be compared exactly, such as numbers or labels, and it cannot fix errors every path shares.
Self-consistency samples independent reasoning chains and takes a majority vote: . If each chain has accuracy , the majority vote accuracy is , which converges to 1 as by the law of large numbers. For and : majority vote accuracy rises to ~83%. The key assumption is that errors are independent across chains — using temperature sampling () ensures diversity in reasoning paths, approximating independence.
On a math problem, 5 CoT samples produce answers: [42, 42, 38, 42, 38]. What is the self-consistency answer?
Theory Exercise
Problem:
A student claims that CoT prompting is just 'making the model show its work' and doesn't actually improve reasoning ability. Argue for or against this claim, citing evidence.
Hints:
- Think about what happens computationally with and without CoT
- Consider the effective compute time per problem
- Think about whether CoT changes what the model can compute
Related Problems on PixelBank
A production assistant talks to thousands of users, and each of them starts a conversation without knowing its rules. The product still needs every one of those conversations to follow the same persona, stay in scope, respect the same safety limits, and format answers the same way. Repeating those rules in every user message is impossible, because users write the user messages.
The previous topic, Chain-of-Thought, shaped how a model reasons on one task. This topic shapes how it behaves for a whole conversation, with the system prompt, the developer-written message that comes before any user turn.
We start with system prompt structure: role, behaviour, constraints, output format, and domain knowledge. Then we cover persona engineering, which tunes vocabulary and depth to the audience. Next come guardrails and safety, including why prompt injection means a system prompt is a policy statement, not a security boundary. We finish with system prompt optimisation, where a test suite of representative inputs turns prompt edits into measured changes.
Definition
A system prompt is a developer-authored message placed before the conversation, in a separate system role of the chat template. It defines the assistant's role, behaviour, constraints, and output format for every turn. Instruction-tuned models are trained to give it priority over user messages when the two conflict, but that priority is learned, not enforced, so injected text can still override it.
In this topic
System Prompt Structure
A good system prompt reads like a brief for a new employee. The role and identity say who the assistant is and for which product. Behaviour guidelines set tone and approach. Constraints list what is out of scope and what must never happen. The output format fixes structure, length, and markup. Knowledge context supplies facts the model cannot know, such as product names or policies, and optional examples show the desired behaviour. Put the most important rules first and group them under headings. The failure mode is a long list of unrelated rules added over time, which contradict each other and dilute the ones that matter.
The system prompt occupies the first position in the context window: . Due to the causal attention mask, every generated token attends to the system prompt, making it a persistent bias on the entire output distribution. The influence of the system prompt on token is approximately — the attention weight from position to the system prompt tokens. In early layers, is high (system prompt provides the behavioral context); in later layers, recent tokens dominate.
Write a system prompt for a customer support bot for a SaaS product.
Persona Engineering
A persona in the system prompt sets the assistant's vocabulary, level of detail, tone, and what it assumes the user knows, and it stays in force for the whole conversation. The same model can act as a rigorous professor or a patient tutor for children, depending only on this description. A good persona is concrete: it names the audience, gives example behaviours, and says what to do when the user is stuck. It must match the product, because a playful persona in a legal tool erodes trust. Personas change style, not knowledge, and a long conversation can drift away from them, so test later turns as well.
Persona engineering conditions the model on a role description that activates a specific region of the learned output distribution. Formally, the persona defines a subspace of outputs where the model concentrates probability mass: . The effectiveness scales with specificity: 'You are a helpful assistant' provides weak conditioning (broad ), while 'You are a board-certified oncologist specializing in lung cancer immunotherapy' provides strong conditioning (narrow ). The cross-entropy between persona-conditioned and unconditioned distributions measures persona strength: .
Design two different personas for the same math tutoring task: one for university students, one for 10-year-olds.
Guardrails and Safety
Guardrails in the system prompt state what the assistant must refuse, which data it must not reveal, what is out of scope, and when to hand off to a person. They reduce harmful or off-topic output in normal use. They are not a security boundary. Prompt injection, first described publicly in 2022, lets user text or retrieved content such as a web page override instructions, because the model reads all of it as one token stream. Training models on an instruction hierarchy that ranks system above user above tool output improves robustness but does not guarantee it. Enforce hard rules in code outside the model.
Safety guardrails define a rejection region where the model should refuse to comply. The challenge is that is defined in natural language (e.g., 'harmful content') and has fuzzy boundaries. In the model's representation space, the decision boundary between compliance and refusal is a hyperplane where is the hidden representation. Adversarial attacks find inputs near this boundary: where is small but changes sign. The robustness of guardrails depends on the margin .
A financial advisor bot should never provide specific investment advice (regulatory requirement). How do you enforce this?
System Prompt Optimization
System prompts are improved empirically, like code. Start with a minimal prompt and run it on a fixed test suite of representative inputs, including edge cases and adversarial ones. Collect the failures, add one targeted instruction for each, and re-run the whole suite, because a new rule often breaks an old behaviour. Keep prompts under version control with their test results, so every change is measurable and reversible. Remove instructions that no longer change any test outcome, since the system prompt is paid for on every request. At 800 tokens and 50,000 requests a day, that is 1.2 billion input tokens a month.
System prompt optimization searches over the space of natural language instructions to maximize a quality metric where is the system prompt, is the model's output given prompt and system prompt , and score measures quality. This is a black-box optimization problem in a discrete (text) space. Methods: APE (Automatic Prompt Engineer) uses the LLM itself to propose prompt candidates, evaluates them on a validation set, and iterates. The search space is exponential in prompt length, so heuristic methods (beam search over prompt variations, evolutionary strategies) are necessary.
Your chatbot sometimes responds in the wrong language (user writes in English, bot responds in French). How do you fix the system prompt?
Theory Exercise
Problem:
Design a comprehensive system prompt for an AI coding assistant integrated into an IDE. The assistant should help with code completion, debugging, and explanation. It should be aware of the user's project context. What sections would you include?
Hints:
- Think about what makes a great coding assistant
- Consider the IDE context (current file, language, project structure)
- Think about safety (not executing arbitrary code, not leaking secrets)
When an LLM becomes part of a software system, its reader is often a program, not a person. A JSON parser does not forgive a missing brace, a trailing comma, a friendly sentence before the object, or a number written as a word. One malformed response in a thousand is enough to break a pipeline that runs a million times a day. So we need output that is not just correct in meaning but valid by construction.
The previous topic, System Prompts, set rules for a whole conversation. This topic sets rules for the exact shape of each answer, so that code can consume it directly.
We start with JSON mode, the simplest way to request machine-readable output, and what it does and does not guarantee. Then we cover schema specification, describing the fields, types, and constraints with a JSON Schema, TypeScript type, or Pydantic model. Next comes constrained decoding, which masks invalid tokens so the output must follow a grammar. We finish with validation and error handling: checking every output, retrying with the error message, and falling back safely.
Definition
Structured output is model output that must conform to a machine-readable format, such as JSON matching a given schema, so a program can parse it without human review. It is obtained through prompting with a schema, provider JSON modes, or constrained decoding that masks tokens a grammar forbids. Only constrained decoding guarantees syntactic validity, and none guarantees the values are correct.
In this topic
JSON Mode
JSON mode is a provider setting that makes the model return syntactically valid JSON, usually by constraining decoding. It guarantees the output parses, but not that it has the fields you need, so it should be paired with a schema. Without native support, you can get reliable JSON by giving the exact schema in the prompt, adding one or two examples, asking for JSON only with no surrounding prose, and ending the prompt where the object should begin. Post-processing can repair common slips such as code fences or trailing commas. The failure mode is JSON that parses but has missing fields or wrong types.
JSON mode constrains the model's output to valid JSON by modifying the sampling process at each step: . This is context-free grammar (CFG) guided decoding where the grammar is the JSON specification. At each token position, only tokens that maintain valid JSON are allowed — the valid set depends on the parser state (e.g., after '"key":', only '"', number, '{', '[', true, false, null are valid starts). The constraint never changes the relative probabilities of valid tokens — only invalid tokens are zeroed out.
Design a prompt that extracts event information as JSON: {"name": str, "date": str, "location": str, "attendees": int}.
Schema Specification
A schema states the expected fields, their types, which are required, and constraints such as allowed values, ranges, and maximum lengths. Models have seen JSON Schema, TypeScript interfaces, and Pydantic classes throughout their training data, so these formats are understood with little explanation. A formal schema removes ambiguity that prose leaves, such as whether a price is a number or a string with a currency sign. The same schema can then validate the output in code, so the prompt and the check never disagree. Comments in the schema, such as the scale of a rating, carry the meaning types cannot.
JSON Schema defines a type system over the output space: each field has a type (, , , etc.), constraints (, , ), and optionality. The schema partitions the output space into valid and invalid sets. The model's effective output distribution becomes for . The renormalization factor measures how much the schema constrains the output: if , the model must concentrate 100 more probability mass on valid outputs.
Use TypeScript types to specify the output format for a product extraction task.
Constrained Decoding
Constrained decoding enforces the format during generation. At each step a grammar or schema determines which tokens could legally come next. The formula sets every other token's probability to zero and renormalises the rest, dividing by the total probability of the valid set. Outlines compiles a regular expression or JSON Schema into a finite-state machine, so the valid set is found in near-constant time per token. Output is then valid by construction. The risk is that forcing a format can push the model into lower-probability text, and Tam et al. (2024) reported reasoning losses under strict format constraints, which later studies disputed.
Constrained decoding implements a finite-state machine (FSM) or pushdown automaton (PDA) that tracks the parser state during generation. At state , the set of valid next tokens is . The modified distribution is where . For JSON, the PDA has states for: object-start, key, colon, value, comma, array, nested structures. The constraint eliminates ~99%+ of the vocabulary at each step (e.g., after '{', only '"' and '}' are valid), dramatically reducing effective perplexity.
After generating '{"name": "John", "age": ', the valid next tokens for constrained JSON generation are what?
Error Handling and Validation
Every structured output should be validated before use, even with JSON mode or constrained decoding, because a schema-valid object can still contain wrong values. Parse it, validate it against the schema, and check business rules such as dates in the past. On failure, retry by sending the output back with the exact validation error, which models usually fix in one attempt. Cap the retries, then fall back to a safe default or flag the item for a person. If each attempt is valid with probability 0.95, two independent failures happen only 0.25 percent of the time. Log every failure, because those logs show which prompt instructions need work.
Even with constrained decoding, semantic errors can occur: the JSON is syntactically valid but semantically wrong (e.g., age = -5, email = 'not-an-email'). Post-generation validation applies the full schema including semantic constraints: . Retry strategies on validation failure use the error message as feedback: , giving the model a chance to self-correct. The expected number of retries before success is where is the per-attempt success probability — for well-prompted models, and rarely more than 2 attempts are needed.
The model returns: {"name": "John", "age": "thirty"}. The schema requires age as a number. How do you handle this?
Theory Exercise
Problem:
You are building a pipeline where an LLM extracts information from emails, produces JSON, and the JSON is processed by a downstream system. The downstream system crashes on invalid JSON. Design a robust pipeline that never passes invalid data downstream.
Hints:
- Think about multiple layers of validation
- Consider retry strategies
- Think about what happens when retries also fail