PIXELBANKv9.1.0
Menu
Back to LLM Study Plan
Week 7-8

Chapter 7: Prompt Engineering

Master the art and science of communicating effectively with LLMs. Learn zero-shot and few-shot prompting techniques, Chain-of-Thought reasoning that unlocks step-by-step problem solving, system prompts for controlling model behavior, and structured output techniques for reliable integration with downstream systems.

Chapter Overview

Prompt engineering is the practice of designing inputs that elicit desired outputs from LLMs. It is the primary interface between human intent and model behavior---no code changes, no fine-tuning, just carefully crafted text. Despite its apparent simplicity, effective prompting can dramatically improve model performance on complex tasks.

The core insight of prompt engineering is that LLMs are sensitive to how questions are framed. The same question asked differently can produce answers that vary from incorrect to expert-level. This sensitivity arises because the model's behavior is shaped by patterns in its training data: it responds differently to a casual question than to a formal academic query, differently to a vague instruction than to a precise one.

Modern prompt engineering has evolved from simple instruction-writing to systematic techniques like Chain-of-Thought (CoT) reasoning, which can improve math accuracy by 50%+, and structured output prompting, which enables reliable integration with software systems. Understanding these techniques is essential for any LLM practitioner.

This chapter covers:

  • Zero-Shot Prompting: Getting results without any examples
  • Few-Shot Prompting: Teaching by example with in-context demonstrations
  • Chain-of-Thought: Unlocking step-by-step reasoning for complex problems
  • System Prompts: Configuring model behavior, persona, and constraints
  • Structured Output: Generating reliable JSON, XML, and formatted data

Chapter Roadmap

Click any topic to jump in

1
Zero-Shot

Clear instructions, role assignment, and task decomposition — getting results without any examples.

Clear InstructionsRole AssignmentTask Decomposition in PromptsNegative Instructions
2
Few-Shot

In-context learning with demonstrations — how example selection and formatting drive model behavior.

In-Context Learning (ICL)Shot Count and QualityExample Selection StrategyFormatting Consistency
Adding reasoning to prompts
3
Chain-of-Thought

Step-by-step reasoning that unlocks math and logic — zero-shot CoT, few-shot CoT, and self-consistency.

Why CoT WorksZero-Shot CoTFew-Shot CoTSelf-Consistency
Production prompting patterns

Control and reliability for deployed systems

4
System Prompts

Persona engineering, guardrails, and behavioral constraints — configuring model behavior at the system level.

System Prompt StructurePersona EngineeringGuardrails and SafetySystem Prompt Optimization
5
Structured Output

JSON mode, schema specification, and constrained decoding — reliable machine-readable output from LLMs.

JSON ModeSchema SpecificationConstrained DecodingError Handling and Validation

A trained, aligned model still does nothing useful until someone asks it for something, and the way you ask changes what you get. The same model can return a vague essay or a precise, correctly formatted answer depending only on the words in the prompt. No weights change between those two outcomes. So the first skill in using an LLM is writing a request that leaves the model as little room to guess as possible.

The previous chapter, RLHF & Alignment, trained models to follow instructions and prefer helpful answers. This chapter, Prompt Engineering, is about using that ability well, and this first topic covers the simplest case: a bare instruction with no examples.

We start with clear instructions, the habit of stating the task, audience, length, and format explicitly. Then we cover role assignment, where a persona shifts the vocabulary and depth of the answer. Next comes task decomposition, which spells out the steps a complex request needs. We finish with negative instructions, which rule out the failure modes you have already seen, and with why those work best when paired with a positive alternative.

Definition

Zero-shot prompting asks a model to perform a task from a natural-language instruction alone, with no worked input-output examples in the prompt. The model must infer the task, the expected format, and the level of detail from the instruction and its pretrained and instruction-tuned knowledge. Output quality therefore depends directly on how specific and unambiguous the instruction is.

In this topic

1Clear Instructions
2Role Assignment
3Task Decomposition in Prompts
4Negative Instructions
1 of 4
Clear Instructions

A model reads a prompt as the start of a document and continues it in the most likely way. A vague request such as "summarize this" is compatible with thousands of continuations, so the model picks a generic middle ground. Every detail you add removes candidates: the audience, the length, the format, the focus, and what counts as done. Good instructions name the task, the input, the constraints, and the output shape, in that order. Put long input text in clearly delimited blocks so it is not confused with the instruction. The failure mode is under-specification, and the fix is to write the prompt a careful new colleague would need.

Mathematical Intuition

Prompt specificity reduces output entropy: a vague prompt xvx_v gives H(Y∣xv)≫H(Y∣xs)H(Y \mid x_v) \gg H(Y \mid x_s) for a specific prompt xsx_s, where HH is the conditional entropy of the model's output distribution. Each constraint in the prompt eliminates a fraction of the output space: with kk independent constraints each reducing options by half, the remaining space is 2−k2^{-k} of the original. In practice, constraints are correlated, so the reduction follows ∣Yconstrained∣≈∣Y∣⋅e−αk|\mathcal{Y}_{\text{constrained}}| \approx |\mathcal{Y}| \cdot e^{-\alpha k} where α<ln⁡2\alpha < \ln 2 is the average constraint strength.

Example:

Compare the outputs for: (A) "Tell me about dogs" vs (B) "List 5 key facts about domestic dog care that a first-time owner should know, in bullet points."

2 of 4
Role Assignment

A role, such as "You are an emergency physician", tells the model which part of its training distribution to imitate. Text written by specialists uses different vocabulary, assumes different background knowledge, and covers different risks than text written for a general audience. The role shifts all of those at once, which is why one sentence can change the depth of an answer. Roles work best when they come with the audience and the goal, not alone. They do not add knowledge the model lacks, and research on factual benchmarks finds persona gains small and inconsistent. So use roles to set register and depth, not to raise accuracy.

Mathematical Intuition

Role assignment shifts the model's conditional distribution from P(y∣x)P(y \mid x) to P(y∣role,x)P(y \mid \text{role}, x). By Bayes' rule, this is proportional to P(role∣y,x)⋅P(y∣x)P(\text{role} \mid y, x) \cdot P(y \mid x) — the model up-weights responses that are consistent with the assigned role. The effectiveness depends on how well the role is represented in pretraining data: 'expert cardiologist' activates patterns from medical texts (high-quality training signal), while 'alien from planet Zorg' has no grounding and mainly affects style. The expected quality improvement is proportional to DKL(P(y∣role,x)∥P(y∣x))D_{\text{KL}}(P(y \mid \text{role}, x) \| P(y \mid x)) — larger distributional shifts indicate the role is having more effect.

Example:

Ask "What causes chest pain?" with (A) no role and (B) "You are an emergency medicine physician."

3 of 4
Task Decomposition in Prompts

A single instruction such as "analyse this data" hides several sub-tasks, and the model may do only the first one or blend them together. Decomposition lists the steps explicitly, in the order they should run: identify trends, then anomalies, then causes, then recommendations. Numbered steps act as a checklist, so each part of the output maps to one requirement, and a missing section is easy to spot. It also sets the order of reasoning, so later steps can use earlier results. The cost is prompt length and some rigidity. For tasks with dependent stages, chaining separate calls, each with one step, gives even more control.

Mathematical Intuition

Decomposing a task into kk sequential steps reduces the effective complexity from O(Ck)O(C^k) to O(k⋅C)O(k \cdot C) where CC is the complexity per step. This is because the model processes one step at a time, each with bounded context. The error probability also changes: for independent steps with per-step error ϵ\epsilon, the overall success probability is (1−ϵ)k≈1−kϵ(1 - \epsilon)^k \approx 1 - k\epsilon for small ϵ\epsilon. With k=5k = 5 steps and ϵ=0.05\epsilon = 0.05 per step, overall success is ~77%. Without decomposition, the single-step error for the full task might be 50%+ because the model must implicitly manage all subtasks simultaneously.

Example:

You need a model to review a pull request. Write a zero-shot prompt with task decomposition.

4 of 4
Negative Instructions

Negative instructions tell the model what to avoid: no disclaimers, no jargon, no invented facts. They are the natural fix once you have seen a specific failure. They work best when paired with a positive alternative, such as "say I don't know if the context does not contain the answer". A bare prohibition leaves the model without a replacement behaviour, and naming an unwanted phrase can even prime it. Prohibitions are also soft: a long conversation or a conflicting user request can override them. For hard rules, such as never outputting personal data, add a check in code, not only a prompt instruction.

Mathematical Intuition

Negative instructions ('do not X') constrain the output distribution by zeroing out probability mass on undesired outputs: P′(y)∝P(y)⋅1[y∉Yforbidden]P'(y) \propto P(y) \cdot \mathbb{1}[y \notin \mathcal{Y}_{\text{forbidden}}]. The effectiveness depends on how much probability mass the model originally placed on forbidden outputs. If P(y∈Yforbidden)=pP(y \in \mathcal{Y}_{\text{forbidden}}) = p, then after the constraint, the remaining outputs are renormalized by 1/(1−p)1/(1-p). A common failure mode: the constraint is too vague ("don't be verbose") so Yforbidden\mathcal{Y}_{\text{forbidden}} is not well-defined, and the model cannot reliably exclude those outputs.

Example:

A model keeps adding 'I hope this helps!' to every response. How do you fix this?

Theory Exercise

Problem:

You need an LLM to extract key information from job postings. Design a zero-shot prompt that extracts: job title, company, location, salary range, required skills, and experience level. The output should be structured.

Hints:
  • Think about the output format you need
  • Consider edge cases (salary not listed, remote work, etc.)
  • Be explicit about what to do when information is missing