Chapter 3: Self-Improvement: Reflexion, Self-Refine, Tree-of-Thought, LATS
Agents that learn from their own mistakes — without weight updates. Reflexion's verbal reinforcement learning, Self-Refine's critic loops, Tree of Thoughts' deliberate search over reasoning paths, and Language Agent Tree Search (LATS) which combines them all.
- Topics
- 5
- Demos
- 2
- Min read
- 15
Chapter Overview
A bare ReAct agent that fails has no way to recover. The next time it sees a similar problem, it makes the same mistake. Self-improvement techniques fix this without retraining the model — they let the agent inspect its own trajectory, write a critique, and use the critique as additional context on the next attempt.
Four techniques dominate:
- Reflexion (Shinn et al., 2023) — verbal reinforcement learning. After a failed attempt, the model writes a paragraph of self-critique into an episodic memory. The next attempt sees the critique in its prompt.
- Self-Refine (Madaan et al., 2023) — single-shot critique-and-revise. The model produces an answer, then critiques its own answer, then revises based on the critique. Iterate until convergence.
- Tree of Thoughts (Yao et al., 2023) — explore the tree of reasoning paths instead of committing to a single chain. Use BFS/DFS with an LLM-as-judge to score branches.
- Language Agent Tree Search (LATS) (Zhou et al., 2023) — unify ReAct, Reflexion, and ToT under MCTS. The state-of-the-art for hard reasoning + acting tasks.
The meta-insight: for many tasks, more inference-time compute on a smaller model beats more training-time compute on a larger model. Self-improvement is the most effective way to spend that inference compute.
This chapter covers:
- Reflexion: Verbal RL — the loop of attempt → reflect → retry
- Self-Refine & Critic loops — generate, critique, revise
- Tree of Thoughts — search over reasoning paths with LLM-as-judge
- LATS — MCTS over agent trajectories
- Actor / Critic / Reflection tradeoffs — what each technique buys you, what it costs
Chapter Roadmap
Click any topic to jump in
Reflexion (Verbal RL)
Write a paragraph of self-critique, feed it into the next attempt. Gradient-free policy improvement.
Same trajectory (Self-Refine) or search the tree (ToT)
Self-Refine
Generate, critique, revise — single-shot self-improvement for one-off generation tasks.
Tree of Thoughts
Search the tree of reasoning paths with LLM-as-judge scoring.
LATS (MCTS Unifier)
MCTS over agent trajectories — combines ReAct, Reflexion, and ToT into one framework.
Where to Spend Compute
A priority order for inference-time techniques when your agent is failing.
Topics
5 topics in this chapter
A bare ReAct agent that fails has no memory of the failure: face it with a similar problem tomorrow and it repeats the same mistake, because nothing in its loop carries a lesson from one attempt to the next. Chapter 2 gave you better plans, but a good plan still fails against a surprising world. Reflexion (Shinn et al., 2023) is the first and simplest self-improvement technique, and it fixes recovery without touching a single weight. After a failed attempt the agent writes a paragraph of natural-language self-critique — an episodic memory of what went wrong — and that paragraph is prepended to the next attempt's prompt. This topic builds the three-role loop of Actor, Evaluator, and Self-Reflector that makes it work, and explains why it is called verbal reinforcement learning: it is policy improvement carried out in prompt space rather than parameter space. The reason it works at all is that strong LLMs learn from specific textual feedback far more readily than they plan a flawless trajectory on the first try — a theme that runs through every technique in the chapter.
Definition
Reflexion is a self-improvement loop in which an Actor produces a trajectory, an Evaluator scores it, and a Self-Reflector writes a natural-language critique of any failure into episodic memory, which is fed into the next attempt's prompt. It improves the policy through in-context textual feedback rather than weight updates — verbal reinforcement learning.
In this topic
- 1The Three-Component Loop
- 2Why Verbal RL Beats Numerical RL Here
The Three-Component Loop
Reflexion factors the agent into three roles, each played by an LLM call:
- Actor. A standard ReAct agent that takes actions in the environment. Outputs a trajectory.
- Evaluator. Examines the trajectory and emits a binary or scalar score (did the task succeed? how well?). Can be a programmatic check (test passed?), an LLM-as-judge call, or a human.
- Self-Reflector. Given a failed trajectory and the evaluator's score, writes a paragraph of natural-language reflection: "I failed because I forgot to escape the SQL string. Next time, sanitize all user input before query construction."
The reflection is appended to an episodic memory buffer. On the next attempt, the actor's prompt includes the most recent reflections (typically the last ). The actor reads the reflections and adjusts its strategy.
Reflexion's improvement can be framed as policy improvement under a textual representation. The reflection is a hint sampled from — the conditional distribution over hints given the failed trajectory. The next attempt's policy is . This is policy iteration in prompt space rather than parameter space — gradient-free, but it works because the in-context conditioning of strong LLMs is approximately as expressive as a small fine-tune.
An agent attempts a coding task and fails the test suite. What does Reflexion add to the next attempt?
Why Verbal RL Beats Numerical RL Here
Classical RL gives the policy a scalar reward and updates parameters via backprop. For agentic tasks this is hard:
- The reward signal is sparse (only at the end of the trajectory)
- Credit assignment across 50+ steps is brittle
- Off-policy data is expensive to collect
Verbal RL sidesteps all of this. The reflection IS the credit assignment — written in natural language by an LLM that is much better at counterfactual reasoning ('what if I had done X instead?') than a value function is at backing up scalars. And because the reflection is text, it can be highly specific to the task: 'don't forget to escape SQL strings' is a much more useful signal than 'reward = -0.3'.
Related Problems on PixelBank
Reflexion improves an agent across whole trajectories, but many valuable tasks are single-shot generations — write this function, summarize this document, draft this SQL — where there is no multi-step trajectory to reflect on. Self-Refine (Madaan et al., 2023) is Reflexion's single-shot cousin, and it is the workhorse of in-context self-improvement for exactly these one-off tasks. The same model generates an output, then critiques its own output, then revises based on that critique, looping until the critique reports nothing left to fix or a step budget runs out. This topic builds that generate/critique/revise cycle and explains the slightly surprising reason it helps: the critique prompt puts the model into an adversarial cognitive mode distinct from the constructive mode that generated the text, so the same weights can find errors they just produced. It also draws the sharp lines between three easily-confused cousins — Self-Refine, Reflexion, and LLM-as-Judge — because production systems routinely nest them (a Reflexion evaluator is often itself an LLM-as-Judge call), and knowing which is which keeps that composition clear.
Definition
Self-Refine is a single-model, single-trajectory improvement loop that generates an output, critiques it, and revises based on the critique, iterating until convergence or a step budget. Unlike Reflexion, which improves across multi-step trajectories, Self-Refine improves within one user-visible response, exploiting the model's distinct generation and critique modes.
In this topic
- 1The Generate / Critique / Revise Cycle
- 2Self-Refine vs Reflexion vs LLM-as-Judge
The Generate / Critique / Revise Cycle
Three prompts, one model:
- Generate. Standard one-shot generation: 'Write a Python function that returns the -th Fibonacci number.'
- Critique. Same model, new prompt: 'Here is a function: <output>. Identify any bugs, inefficiencies, or missing edge cases.'
- Revise. Same model, third prompt: 'Here is the function and a critique. Produce an improved version.'
Iterate steps 2-3 until either (a) the critique reports 'no further changes needed' or (b) you hit a max iteration count. Empirically 2-4 iterations suffices.
Why it works: the critique prompt invokes a different cognitive mode than the generation prompt. The generator is constructive; the critic is adversarial. The same model in critic mode can spot errors its generator mode produced — because generation and critique pull on different parts of the prompt distribution.
Self-Refine vs Reflexion vs LLM-as-Judge
Three close cousins, distinct uses:
- Self-Refine. Same model, single trajectory. Used for one-shot tasks (write, summarize, code). Improvement happens within a single user-visible response.
- Reflexion. Same model, multi-trial. Used for multi-step agent tasks. Improvement happens across trajectories; the actor sees critiques from prior trials.
- LLM-as-Judge. Different model (usually larger), single trajectory. Used for evaluation, not improvement. The judge model scores the actor model's output for benchmarks or RLHF reward.
In production you often combine them: an actor agent does multi-trial Reflexion; the evaluator inside Reflexion is itself an LLM-as-Judge call.
You are building a system that generates SQL from natural language. Where does each technique apply?
Chain-of-thought reasoning, and the Reflexion and Self-Refine loops built on it, all commit to a single line of reasoning and try to make it good. But on problems where the right first step is non-obvious — Game of 24, puzzles, constrained creative writing — a single chain that starts down the wrong path rarely recovers, and no amount of critique of that one chain finds the path it never took. Tree of Thoughts (ToT) (Yao et al., 2023) generalizes the single chain into a tree of branching reasoning paths. At each node the model proposes several candidate next thoughts, an LLM-as-judge scores how promising each partial path is, and a search algorithm — BFS or DFS with pruning — explores the tree, reading the answer from the best-scoring leaf. This topic builds the four components (state, generator, evaluator, search) and quantifies the trade: on Game of 24, single-chain CoT solves about 4% while ToT solves roughly 74%, at close to 10× the inference cost. That cost is the whole story of when ToT is worth it, and it sets up LATS next.
Definition
Tree of Thoughts (ToT) generalizes chain-of-thought into a search over a tree of partial reasoning paths: at each node the model proposes several candidate next thoughts, an evaluator scores each partial path, and BFS or DFS with pruning explores promising branches. The answer is read from the highest-scoring leaf, trading large inference cost for far higher success on hard search problems.
In this topic
- 1Branching, Scoring, Pruning
- 2When ToT Pays for Itself
Branching, Scoring, Pruning
ToT requires four components:
- State representation. A 'thought' is a partial reasoning step. The state is the path from root to the current node.
- Thought generator. Given a state, produce candidate next thoughts (e.g., different sub-strategies).
- State evaluator. An LLM-as-judge or heuristic that scores how 'promising' a state is (probability of leading to a correct answer). Can be: 'rate 1-10' or 'sure / likely / impossible'.
- Search algorithm. BFS (good for fixed-depth puzzles) or DFS (good for open-ended exploration). Beam search keeps the top- at each level.
The Game of 24 example from the original paper: given 4 numbers, reach 24 using +, -, ×, /. CoT alone solves ~4%. ToT (BFS, , ) solves ~74%.
When ToT Pays for Itself
ToT is expensive — at branching factor 3 and depth 4, you do up to generation+scoring calls per problem. It pays off when:
- The problem has a known short solution but the right path is non-obvious (the model often picks a wrong first step)
- You can afford the cost (single high-stakes decision, batch processing)
- A successful judge exists (the LLM-as-judge can reliably score partial states)
It does NOT pay off when the problem requires real-world tool calls in the loop (each branch would need to execute tools, blowing up the cost) or when the solution requires long paths (depth 10+) where the search budget runs out.
Related Problems on PixelBank
The techniques so far each address one weakness in isolation: ReAct acts but cannot search plans, Tree of Thoughts searches reasoning but takes no real actions, and Reflexion recovers from failure but explores only one trajectory at a time. Hard tasks fail along all three axes at once, which is what LATS (Language Agent Tree Search, Zhou et al., 2023) is built to handle. It is the synthesis of the whole chapter: take Tree of Thoughts' search structure, ReAct's real tool-using trajectories, and Reflexion's self-critique, and unify them under Monte Carlo Tree Search. This topic builds that MCTS loop over agent trajectories — select by UCB1, expand with candidate actions, simulate a rollout, reflect on failures, backpropagate the outcome — and shows why addressing bad plans, bad execution, and failed recovery simultaneously is what let LATS reach state-of-the-art on HotpotQA, ALFWorld, and SWE-bench Lite. It closes with the unavoidable caveat: LATS costs orders of magnitude more LLM calls than plain ReAct, so it belongs to single high-stakes tasks, not high-volume serving.
Definition
LATS (Language Agent Tree Search) unifies ReAct's tool-using trajectories, Tree of Thoughts' search, and Reflexion's self-critique under Monte Carlo Tree Search. It selects paths by UCB1, expands with candidate actions, simulates rollouts, injects reflections on failure, and backpropagates outcome scores — addressing bad planning, bad execution, and failed recovery together at very high inference cost.
In this topic
- 1MCTS over Agent Trajectories
- 2Why LATS Wins on Hard Tasks
MCTS over Agent Trajectories
LATS treats each ReAct trajectory as a path in a tree. Nodes are agent states; edges are (thought, action, observation) triples. Standard MCTS:
- Select. From the root, descend along edges chosen by UCB1 — balancing the highest-value path with under-explored siblings.
- Expand. At a leaf, sample candidate next thoughts and execute the corresponding actions.
- Simulate. Continue rollout for several steps using greedy ReAct.
- Reflect. If the rollout failed, write a reflection (Reflexion-style) and inject it back as additional context.
- Backpropagate. Update the value estimate of every node along the selected path with the rollout's outcome score.
After MCTS iterations, the highest-value root-to-leaf path is the final trajectory.
LATS state-of-the-art on HotpotQA, ALFWorld, and SWE-bench Lite at the time of release. It is also extremely expensive — orders of magnitude more LLM calls than ReAct alone.
Why LATS Wins on Hard Tasks
Hard tasks fail in three ways: bad plan, bad execution, no recovery. LATS addresses all three:
- Bad plan: MCTS explores many candidate plans (Tree of Thoughts component)
- Bad execution: each branch is a real ReAct trajectory with tool calls (ReAct component)
- No recovery: failures generate reflections that influence subsequent MCTS iterations (Reflexion component)
For a single high-stakes task — a one-off code refactor, a contract review, a research synthesis — LATS is the right tool. For high-volume serving (chatbots, customer support), it is overkill; the per-request cost is prohibitive.
Related Problems on PixelBank
Having met Reflexion, Self-Refine, Tree of Thoughts, and LATS, you can see they are all one idea in different clothing: spend more inference-time compute to make up for an imperfect base model. The engineering question is never 'which is most powerful' but 'how much extra compute is worth it, and where should it go' — and the answers span more than two orders of magnitude, from Self-Refine's roughly 3× to LATS's 100–300×. This closing topic lays the techniques on a single cost frontier and, more usefully, gives a priority order for what to try when an agent is failing. The counterintuitive punchline, which most teams get wrong, is that the cheapest interventions dominate: better tools and observations are free wins that dwarf every prompting trick, a better system prompt costs a day of work and nothing at inference, and only after those are exhausted does it make sense to climb into Self-Refine, Reflexion, and — rarely — the search-based methods. Getting this ordering right is the difference between an agent that improves and a bill that merely grows.
Definition
The actor/critic/reflection trade-off is the cost-quality frontier of inference-time techniques, ordered roughly ReAct (1×), Self-Refine (~3×), Reflexion (~3–4×), Tree of Thoughts (~50–150×), and LATS (~100–300×). Because quality gains are sublinear in cost, the practical priority is to exhaust cheap fixes — better tools, better prompts — before spending on search.
In this topic
- 1The Inference-Time Compute Frontier
- 2Where to Spend Compute First
The Inference-Time Compute Frontier
Order the techniques by inference cost (LLM calls per task):
- ReAct alone: — call it 1×
- Self-Refine (3 iterations): ~3×
- Reflexion (3 trials): ~3-4× the actor cost + reflector cost
- Tree of Thoughts (, ): ~50-150×
- LATS ( iterations): ~100-300×
The quality gain is rarely linear. ToT might give you 2× success rate at 100× cost — the 'price per percentage point' is wildly different from cheaper techniques. Always ask: would a smaller, cheaper model with Self-Refine match a bigger model with single-shot? Often yes.
Where to Spend Compute First
A practical priority order when an agent is failing:
- Better tools / observations. Free wins. Dwarfs every prompting trick.
- Better system prompt. Spend a day, free at inference time.
- Self-Refine. ~3× cost, often closes 30-50% of the gap on generation tasks.
- Reflexion. ~3-4× cost, requires trial structure (you need re-attempts to be meaningful).
- Tree of Thoughts. 50-150× cost. Reserve for high-stakes single-shot reasoning problems.
- LATS. 100-300× cost. Reserve for the very hardest tasks where nothing cheaper works.
Theory Exercise
Problem:
A team has an agent with 70% success rate on a coding task. They are debating: switch to GPT-5 (2× cost), add Reflexion (3× cost), or add Tree of Thoughts (50× cost). They have a budget of 5× current cost. Which delivers the most success-rate gain?
Hints:
- What does each technique buy you in terms of failure modes addressed?
- Are the failures planning errors, execution errors, or recovery errors?
- Could two of the three be combined within budget?