Search-R1
Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Ö. Arık, Dong Wang, Hamed Zamani, Jiawei Han
Read the Paper on arXivPaper Overview
Search-R1 (Jin, Zeng, Yue, Yoon, Arık, Wang, Zamani & Han — UIUC, UMass Amherst and Google Cloud AI Research; COLM 2025) asks whether the DeepSeek-R1 recipe, reinforcement learning from nothing but an outcome reward, can teach a language model to use a search engine as well as to reason.
The motivation is that the two standard ways of combining LLMs with search both leave the model untrained at the thing that matters. RAG retrieves once, with the user's question as the query. Prompted tool use (IRCoT, ReAct) lets the model search mid-thought, but nothing has optimised how it searches, and supervised tool use needs annotated search trajectories that are expensive to produce at scale.
Search-R1 puts the search engine inside the RL rollout. The model reasons in <think> tags, issues queries in <search> tags, receives passages in <information> tags, and commits in <answer> tags, as many times as it chooses within a budget. Two design decisions make this train well: retrieved-token loss masking, so gradients flow only through text the model wrote, and a plain exact-match outcome reward, with no format reward, process reward or reward model.
Across seven QA benchmarks under a strictly controlled setup, Search-R1 lifts Qwen2.5-7B to 0.431 average EM against 0.304 for RAG and Qwen2.5-3B to 0.325 against 0.270, and beats the same RL recipe without search at every model size. Along the way the model learns to call search more often, to chain queries across hops, and to verify its own answers, none of which it was explicitly taught.
Chapter Roadmap
Click any topic to jump in
Learning to Search
RAG retrieves once and prompted tools are never optimised. Can outcome-reward RL teach an LLM when and what to search?
RL with a Search Engine
The search engine becomes part of the environment, and PPO or GRPO optimise trajectories that interleave reasoning and retrieval.
Rollout and prompt
Multi-Turn Rollout
Generate until a search or answer tag appears, retrieve the top-3 passages, append them, and repeat, within an action budget of B = 4.
Minimal Template
One instruction defines the think, search, information and answer tags and nothing else, so learned behaviours come from RL.
Loss and reward
Retrieved Token Masking
Loss only on tokens the LLM wrote. Retrieved passages condition generation but get no gradient: 0.431 vs 0.343 EM.
Outcome-Only Reward
Exact match on the final answer, with no format reward, process reward or neural reward model.
PPO vs GRPO
GRPO converges faster but can collapse. PPO is slower but stable, and final rewards are comparable.
Evidence and behaviour
Results
0.431 avg EM on 7B vs 0.304 for RAG across seven QA benchmarks, beating RL without search at every size.
Emergent Search
Response length dips, then climbs as the model learns to search more. Top-3 retrieval works best.
Case Studies
Chained queries and self-verification, plus honest failures: span errors, poor decomposition, misleading passages.
A language model's knowledge stops at its training cutoff and is patchy on anything long-tail, so a question like "Curious is a women's fragrance by a singer born in what city and state?" is a trap. It needs two facts chained together, and the model is unlikely to hold both reliably. The obvious fix is to give it a search engine. How it should use that engine is the hard part.
Before Search-R1 there were two families of answer. Retrieval-augmented generation (RAG) runs one retrieval using the user's question as the query and pastes the passages into context. That is cheap and robust, but the query is fixed before the model has thought about anything, so it suits single-hop questions and can fall apart on multi-hop ones, where the second query depends on what the first one returned. Search-as-a-tool methods let the model issue its own queries mid-reasoning. IRCoT and ReAct do this with prompting, which the paper calls suboptimal because the LLM was never optimised to interact with a search engine. Toolformer does it with supervised fine-tuning, which needs large sets of high-quality annotated search trajectories that are expensive to produce at scale. And search is non-differentiable, so you can't simply backpropagate through it.
Reinforcement learning sidesteps both problems. DeepSeek-R1 showed that RL from outcome rewards alone can teach a model to reason, with no annotated chains of thought. Search-R1 asks whether the same recipe can teach a model when and what to search. The authors identify three challenges: (1) framework and stability — how to put a search engine inside an RL loop when part of the trajectory is text the model did not write; (2) multi-turn interleaving — the model should decide dynamically how many searches a question needs; and (3) reward design — whether a simple exact-match outcome reward is enough to induce meaningful search behaviour.
The paper answers all three, and describes itself as an extension of DeepSeek-R1-Zero from parametric reasoning to retrieval-driven reasoning.
Key Points
RAG retrieves once using the input as the query, so the LLM never chooses what to look up and cannot follow up on what it found
Prompted tool use (IRCoT, ReAct) lets the model search mid-reasoning, but nothing optimises how it searches; prompting-based approaches often struggle to generalise
SFT tool use (Toolformer) needs large-scale, high-quality annotated trajectories, and the search call itself is non-differentiable
RL with outcome rewards, as in DeepSeek-R1, needs neither annotated trajectories nor gradients through the search engine
Three open challenges: stable RL with retrieved context, multi-turn interleaved search, and whether a simple outcome reward is sufficient
Search-R1 = DeepSeek-R1-Zero-style RL plus a search engine in the rollout loop
Further Reading
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- IRCoT — Interleaving Retrieval with Chain-of-Thought Reasoning
- ReAct: Synergizing Reasoning and Acting in Language Models
- Toolformer: Language Models Can Teach Themselves to Use Tools
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
- Retrieval-Augmented Generation for Large Language Models: A Survey
Standard RL fine-tuning of an LLM (RLHF and its descendants) maximises a reward on responses sampled from the policy, minus a KL penalty that keeps the policy close to a frozen reference model. It assumes the entire response is written by the policy. Search-R1 breaks that assumption on purpose: during rollout, the model's text is interleaved with passages returned by a search engine .
The paper formalises this by conditioning the policy on the engine: trajectories are drawn from , which it describes as , interleaved retrieval-and-reasoning. The objective keeps the familiar shape (reward minus KL), but both the expectation and the KL term are now taken over these search-augmented trajectories. The search engine is treated as part of the environment. It is frozen, it is never trained, and it only produces observations.
Figure 1 shows the framework is compatible with both mainstream policy-gradient methods. PPO uses a learned value model and Generalized Advantage Estimation to give a per-token advantage. GRPO drops the value model and computes advantages by comparing a group of rollouts for the same question. In both cases the rollout module is the same: a policy LLM that can call the search engine as many times as it decides to, up to a budget. The reward model and reference LLM are frozen.
Treating retrieval as environment feedback rather than as model output is what makes the design work. It is also what forces the next idea, loss masking: if retrieved tokens are observations, the policy-gradient loss should not be computed on them.
Key Points
The search engine is modelled as part of the environment: frozen, never trained, it only returns passages
Trajectories interleave LLM-generated tokens with retrieved passages
The objective is the usual reward minus -weighted KL to a reference policy, with both terms taken over search-augmented trajectories
Works with PPO (actor-critic with GAE) and GRPO (critic-free, group-relative advantages); PPO is the default in the main experiments
Trained models: the policy LLM, plus the value LLM for PPO. Frozen: the reference LLM, the rule-based reward, and the search engine
Hyperparameters (Appendix B): , clip , policy LR 1e-6, value LR 1e-5, GAE , 500 steps on 8×H100
Mathematical Formulation
RL Objective with a Search Engine (Eq. 1)
is a question from dataset , is a full trajectory of reasoning, queries, retrieved passages and answer, is the reward function, and is the frozen reference LLM. Compared with the classical RLHF objective (Eq. 5 in Appendix A), the only change is the conditioning on , but that change means part of was not sampled from the policy at all.
Further Reading
- Proximal Policy Optimization Algorithms
- DeepSeekMath — introduces Group Relative Policy Optimization (GRPO)
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via RL
- HybridFlow / veRL: A Flexible and Efficient RLHF Framework
During rollout the policy is not simply sampling one long completion. Algorithm 1 runs a small generate–detect–retrieve loop. The LLM generates tokens until it emits one of three stop signals: </search>, </answer>, or end-of-sequence. What happens next depends on which one it was.
If the segment contains a <search> query </search> block, the system parses out the query, calls the search engine, and appends the top results to the trajectory wrapped in <information> … </information>. The LLM then keeps generating, now conditioned on what it just retrieved. If the segment contains <answer> … </answer>, the rollout ends and that answer is scored. If the model produced neither, which is a malformed action, the system appends the fixed string "My action is not correct. Let me rethink." and lets it try again.
Each pass through the loop costs one unit of an action budget (set to 4 in the experiments). Search depth therefore isn't fixed by the pipeline. The model decides it, from zero searches on an easy question up to the budget on a hard multi-hop one, and RL rewards whichever choice yields a correct answer. The retriever is E5 over a 2018 Wikipedia dump, returning the top 3 passages per query, with retrieved content capped at 500 tokens.
This is the mechanism behind the paper's case studies: a model that queries "Curious fragrance information", learns it is a Britney Spears perfume, then queries "Britney Spears birthplace". No single-shot RAG query could have chained those two steps.
Key Points
Generation pauses at </search>, </answer> or end-of-sequence, and control passes back to the system
Search detected: parse the query, call , append results inside <information> … </information>, continue generating
Answer detected: return the trajectory for scoring
Neither detected: append "My action is not correct. Let me rethink." and retry
Each turn consumes one unit of the action budget ; the model decides how many of those turns are searches
Retrieval setup: E5 dense retriever, 2018 Wikipedia dump, top-3 passages per query, max 500 tokens of retrieved content
RL can only reinforce behaviour the model sometimes produces, so the starting policy needs some way of expressing a search. Search-R1 supplies it with a single instruction template (Table 1), and very little else.
The template tells the model to reason inside <think> tags every time it gets new information, to call the search engine with <search> query </search> if it lacks knowledge, that results will come back between <information> tags, that it can search as many times as it wants, and to put the final answer inside <answer> tags without detailed illustrations. Then it appends the question.
What it leaves out is the point. The authors limit constraints to structure: the template does not demand reflective reasoning, does not require a search on every question, and does not endorse any problem-solving strategy. That keeps the model's learning dynamics during RL observable and unbiased. Behaviours that emerge later, such as decomposing a multi-hop question, re-querying after a bad result, or running a verification search, are learned from reward and were not prompted.
The same template is used for base and instruct models. A base model barely follows it at first, which is one reason base models start from a lower reward. As later sections show, RL closes that gap.
Key Points
Four tag pairs carry the whole protocol: <think>, <search>, <information> (inserted by the system) and <answer>
The model is told it may search as many times as it wants; nothing forces a search
Structure only: no instructions to reflect, verify, or follow a particular strategy, so learned behaviours can be attributed to RL
The same template is used for base and instruction-tuned Qwen2.5 models
The template is the only scaffold; there is no SFT warm-up on search trajectories before RL
A Search-R1 trajectory is a mix of two kinds of tokens: those the policy generated (thoughts, queries, the answer) and those the search engine returned (the passages inside <information>). Standard PPO and GRPO compute a token-level loss over every token in the rollout. Applied naively here, that would push gradients through the passages too.
That is wrong on two counts. The policy never chose those tokens, so increasing or decreasing their likelihood based on the trajectory's reward assigns credit to text the agent did not produce. It also teaches the model to predict Wikipedia rather than to use it, an objective unrelated to searching well, which the paper describes as unintended learning dynamics.
The fix is a one-line indicator. if token was LLM-generated and if it was retrieved. Both the PPO and GRPO objectives sum only over tokens with and normalise by their count, and in GRPO the KL penalty is masked the same way. The retrieved passages still sit in the context and condition every subsequent token, but they never receive gradient.
The ablation is decisive. On Qwen2.5-7B-base with PPO, masking lifts the seven-dataset average EM from 0.343 to 0.431, and it wins on every dataset. Figure 3 shows higher training reward on both 3B and 7B, and the 3B average rises from 0.262 to 0.303.
Key Points
Rollouts contain policy tokens (think / search / answer) and retrieved tokens (inside <information>)
Unmasked, the policy-gradient loss would update the likelihood of passages the model never chose
marks LLM-generated tokens; losses sum over and divide by
In GRPO the KL-divergence term is also computed only over generated tokens
Retrieved tokens still condition generation; they just don't receive gradient
Qwen2.5-7B-base (PPO): avg EM 0.431 with mask vs 0.343 without. Qwen2.5-3B-base: 0.303 vs 0.262
Mathematical Formulation
Token Loss Mask
A per-token indicator built during rollout. Every token the system inserts inside the information block gets .
Masked PPO Objective (Eq. 2)
Here is the importance ratio and is the GAE advantage computed from future rewards and the learned value function . The clipped surrogate is standard PPO; the only Search-R1 change is that the sum and its normaliser run over generated tokens only.
The training signal is about as simple as it can be. When a rollout finishes, the system extracts the text inside <answer> and checks it against the gold answer with exact match. A match scores 1 and anything else scores 0. Nothing rewards a well-formed query, a relevant retrieval, a sensible number of searches, or the format.
Two omissions are deliberate. First, unlike DeepSeek-R1, there is no format reward. The authors report that the learned model already adheres strongly to the tag structure, and leave richer format rewards to future work. Second, there is no neural reward model. Following DeepSeek-R1, they avoid one because LLMs in large-scale RL are sensitive to the specific form of the reward (a learned reward model can be gamed), and because training one adds compute and complexity.
This is challenge (3) from the introduction: is an outcome reward enough to induce good search behaviour? The rest of the paper says yes. With nothing but EM on the final answer, the model learns to issue more valid searches over training, to decompose multi-hop questions into sequential queries, and sometimes to run a self-verification search before answering. All of that credit assignment flows backward from a single bit at the end of the trajectory.
Key Points
Reward = exact match between the extracted answer and the gold answer; it is binary
No process reward: queries, retrievals and reasoning steps are never scored directly
No format reward, unlike DeepSeek-R1; the model learns the tag structure anyway
No neural reward model, which avoids reward hacking and the cost of training one
Training data: the merged NQ + HotpotQA training sets, giving a mix of single-hop and multi-hop questions
EM is also the evaluation metric across all seven benchmarks
Mathematical Formulation
Rule-Based Outcome Reward (Eq. 4)
is the answer extracted from the answer tags in trajectory and is the ground truth. It returns 1 on an exact string match and 0 otherwise, so one bit at the end of the rollout is the only reward signal.
Search-R1 runs under either algorithm, so the paper compares them directly on Qwen2.5-3B and 7B, base and instruct. GRPO differs from PPO by using the average reward of a group of sampled responses as its baseline, rather than a learned value function: for each question it samples rollouts (5 by default) and scores each one relative to the others.
The paper reports three findings. GRPO converges faster in every setting, because PPO's critic needs warm-up steps before its advantage estimates are useful. PPO is more stable: in Figure 5, GRPO's training reward collapses after many steps on several models, while PPO's does not. And final rewards are comparable, so both are viable. The authors choose PPO as the default because of its stability, and evaluate collapsed runs at their last stable checkpoint.
The evaluation numbers (Table 3) show no single winner. On 7B-base, PPO scores 0.431 average EM against GRPO's 0.350. On 7B-instruct, GRPO leads, 0.396 to 0.385. On 3B, GRPO is ahead for both base (0.312 vs 0.303) and instruct (0.336 vs 0.325).
The group-size ablation (Table 8) sharpens the trade-off. Larger groups converge faster and reach higher training reward but risk collapse. Group size 1, where GRPO reduces to plain REINFORCE, was the most stable and generalised best, at 0.410 average EM against 0.350 for size 5.
Key Points
GRPO converges faster: no critic to warm up
PPO is more stable: GRPO showed reward collapse after extended training
Final training rewards are comparable; PPO is the default for its stability
Table 3 average EM: 7B-base PPO 0.431 vs GRPO 0.350; 7B-instruct GRPO 0.396 vs PPO 0.385
3B: GRPO slightly ahead (base 0.312 vs 0.303; instruct 0.336 vs 0.325)
Group size 1 (= REINFORCE) was the most stable, with 0.410 avg EM vs 0.350 for size 5, so faster learning traded against generalisation
Mathematical Formulation
Masked GRPO Objective (Eq. 3)
The expectation is over and a group , and is the per-token importance ratio. Unlike PPO, the KL term sits directly in the loss rather than in the reward, and it is also restricted to generated tokens.
Group-Relative Advantage (standard GRPO form)
The paper says is computed from the relative rewards of outputs within each group. This is the outcome-supervision form from the original GRPO paper (Shao et al., 2024): every token of rollout gets the same advantage, its EM reward standardised against the group's.
The evaluation covers seven QA benchmarks in two groups: general QA (NQ, TriviaQA, PopQA) and multi-hop QA (HotpotQA, 2WikiMultiHopQA, Musique, Bamboogle). Only NQ and HotpotQA are in-domain, since their training sets were merged for training; the other five are out-of-domain. Every baseline uses the same retriever, the same top-3 passages, the same corpus, the same training data and the same base LLMs, so the comparison isolates the method.
The baselines cover the alternatives: no retrieval (direct inference, CoT); retrieval at inference time (RAG, IRCoT, Search-o1); and fine-tuning (SFT, R1 — the same RL recipe without a search engine — and rejection sampling, which fine-tunes on the policy's own correct search trajectories).
On Qwen2.5-7B, Search-R1-base reaches 0.431 average EM against 0.348 for rejection sampling, the strongest baseline, and 0.304 for RAG, with the largest gains on multi-hop sets (HotpotQA 0.433 vs RAG's 0.299). On Qwen2.5-3B, Search-R1-instruct reaches 0.325 against RAG's 0.270. The paper's headline numbers are not consistent: the abstract claims 24% (7B) and 20% (3B) average relative improvement over RAG baselines, while the contributions list claims 41% and 20%. Computed directly from Table 2 against RAG, the gains are about 42% and 20%. Search-R1 also beats R1 at every size, which isolates the value of search from the value of RL. Appendix C repeats the study on Qwen2.5-14B, where Search-R1-base reaches 0.479, and the gap widens with scale.
The paper's fourth observation is that larger models learn search better: Search-R1's margin over the second-best method is much bigger at 7B than at 3B. One anomaly appears in the 3B-base row, where Bamboogle EM is 0.088, far below the 3B-instruct model's 0.264.
Key Points
7 benchmarks: NQ†, TriviaQA, PopQA, HotpotQA†, 2Wiki, Musique, Bamboogle († = in-domain)
Controlled comparison: same E5 retriever, top-3 passages, 2018 Wikipedia corpus, training data and base LLMs
Qwen2.5-7B: Search-R1-base 0.431 avg EM vs rejection sampling 0.348, RAG 0.304, R1-base 0.276
Qwen2.5-3B: Search-R1-instruct 0.325 vs RAG 0.270, Search-o1 0.255
Qwen2.5-14B (Appendix C): Search-R1-base 0.479 vs R1-base 0.357, Search-o1 0.310
Beats R1 (RL without search) at every size, so the gain comes from search on top of RL
Works for both base and instruct models, and larger models benefit more
Further Reading
- Search-o1: Agentic Search-Enhanced Large Reasoning Models
- IRCoT — Interleaving Retrieval with Chain-of-Thought Reasoning
- HotpotQA: A Dataset for Diverse, Explainable Multi-hop QA
- MuSiQue: Multihop Questions via Single-hop Question Composition
- Measuring and Narrowing the Compositionality Gap (Self-Ask, Bamboogle)
- When Not to Trust Language Models (PopQA)
- Search-R1 (arXiv 2503.09516)
Figure 2 tracks how Search-R1 changes over training, and it describes a learning process rather than a gradual climb.
Response length (Qwen2.5-7B-base, panel c) follows a decrease–increase–stabilise pattern. For roughly the first 100 steps it falls sharply while reward rises only slightly: the base model is dropping filler words and adapting to the task format. After about step 100, both length and reward rise significantly. That is the point where the model starts calling the search engine frequently, and the retrieved passages lengthen the trajectory. Panel d shows the same story from the other side: the number of valid search calls grows steadily over training. Nobody told the model to search more. It learned that searching pays.
Base vs instruct (panel b, Figure 4): instruction-tuned models start higher and converge faster, but final rewards are very similar. General post-training speeds up early learning in this setting, and RL can close the gap for base models over time.
How many passages to retrieve (Appendix G, Table 7) shows a precision–recall trade-off inside RL. Top-5 converges fastest but its reward later declines and becomes unstable; top-1 is stable but limited by recall. Top-3 is best, at 0.431 average EM against 0.375 for top-1 and 0.400 for top-5. The authors suggest that noisy passages at top-5 may teach the model that retrieved context is often unhelpful, discouraging it from relying on search.
Key Points
Early (≈ first 100 steps): response length drops sharply and reward rises slightly as filler words are trimmed
Later (after ≈ 100 steps): length and reward both rise significantly as the model starts searching frequently
Valid search calls increase over training, driven only by the EM reward
Base vs instruct: instruct starts higher and converges faster; the final rewards are very similar
Top-k study: top-3 best (0.431), top-5 (0.400) converges fastest then destabilises, top-1 (0.375) limited by recall
The appendix compares a model trained with RL without search (R1) and Search-R1 on the same question (Table 9): "Curious is a women's fragrance by a singer born in what city and state?" R1 reasons from memory: "The singer is Beyoncé, who was born in Houston, Texas," and answers Houston. It is fluent, confident and wrong.
Search-R1 searches "Curious fragrance information" and learns the perfume is by Britney Spears. It then searches "Britney Spears birthplace" and finds McComb, Mississippi. The paper points to two behaviours here. One is interleaved reasoning and retrieval: each query is written using what the previous one returned. The other is self-verification: after the second search the model already has the answer, yet it issues a third search, "McComb, Mississippi location", to confirm it before answering. This mirrors the self-verification that DeepSeek-R1 observed in pure-reasoning RL.
The paper's eleven additional case studies (Appendix J) are candid about failures, and they map onto the method's limits:
- Answer-span errors under EM. For the Weezer question, the model retrieves the right passage and even names "The Blue Album" in its reasoning, then answers "Weezer".
- Poor decomposition. It sometimes searches the whole question verbatim, repeats an identical query that already failed, and then guesses ("National Magazine Award" instead of Pulitzer Prize).
- Misled by irrelevant retrievals. A passage about a different reality show leads to "London Charles" instead of Mani.
Each of these is a case where one EM bit at the end gives weak guidance, which is why the conclusion names richer reward mechanisms and uncertainty-aware retrieval as next steps.
Key Points
R1 (RL, no search) answers from memory and hallucinates: Beyoncé → Houston
Search-R1 chains two searches, Curious → Britney Spears → McComb, Mississippi
Self-verification: it issues a third, confirming search even after it has the answer
Successful cases include learning to stop searching when the corpus lacks the information, and recovering from a weak first query
Failures: right passage but wrong answer span; failing to decompose and re-issuing identical queries; being misled by irrelevant passages
Future work named in the paper: richer reward mechanisms, uncertainty-driven retrieval, more tools, and multimodal reasoning