PIXELBANKv9.1.0
Menu
Back to Concepts
LLM Agents & Reinforcement Learning2025

Search-R1

Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Bowen Jin, Hansi Zeng, Zhenrui Yue, Jinsung Yoon, Sercan Ö. Arık, Dong Wang, Hamed Zamani, Jiawei Han

Read the Paper on arXiv

Paper Overview

Search-R1 (Jin, Zeng, Yue, Yoon, Arık, Wang, Zamani & Han — UIUC, UMass Amherst and Google Cloud AI Research; COLM 2025) asks whether the DeepSeek-R1 recipe, reinforcement learning from nothing but an outcome reward, can teach a language model to use a search engine as well as to reason.

The motivation is that the two standard ways of combining LLMs with search both leave the model untrained at the thing that matters. RAG retrieves once, with the user's question as the query. Prompted tool use (IRCoT, ReAct) lets the model search mid-thought, but nothing has optimised how it searches, and supervised tool use needs annotated search trajectories that are expensive to produce at scale.

Search-R1 puts the search engine inside the RL rollout. The model reasons in <think> tags, issues queries in <search> tags, receives passages in <information> tags, and commits in <answer> tags, as many times as it chooses within a budget. Two design decisions make this train well: retrieved-token loss masking, so gradients flow only through text the model wrote, and a plain exact-match outcome reward, with no format reward, process reward or reward model.

Across seven QA benchmarks under a strictly controlled setup, Search-R1 lifts Qwen2.5-7B to 0.431 average EM against 0.304 for RAG and Qwen2.5-3B to 0.325 against 0.270, and beats the same RL recipe without search at every model size. Along the way the model learns to call search more often, to chain queries across hops, and to verify its own answers, none of which it was explicitly taught.

Chapter Roadmap

Click any topic to jump in

1
Learning to Search

RAG retrieves once and prompted tools are never optimised. Can outcome-reward RL teach an LLM when and what to search?

framed as
2
RL with a Search Engine

The search engine becomes part of the environment, and PPO or GRPO optimise trajectories that interleave reasoning and retrieval.

realised by

Rollout and prompt

3
Multi-Turn Rollout

Generate until a search or answer tag appears, retrieve the top-3 passages, append them, and repeat, within an action budget of B = 4.

4
Minimal Template

One instruction defines the think, search, information and answer tags and nothing else, so learned behaviours come from RL.

optimised with

Loss and reward

5
Retrieved Token Masking

Loss only on tokens the LLM wrote. Retrieved passages condition generation but get no gradient: 0.431 vs 0.343 EM.

6
Outcome-Only Reward

Exact match on the final answer, with no format reward, process reward or neural reward model.

run under
7
PPO vs GRPO

GRPO converges faster but can collapse. PPO is slower but stable, and final rewards are comparable.

produces

Evidence and behaviour

8
Results

0.431 avg EM on 7B vs 0.304 for RAG across seven QA benchmarks, beating RL without search at every size.

9
Emergent Search

Response length dips, then climbs as the model learns to search more. Top-3 retrieval works best.

10
Case Studies

Chained queries and self-verification, plus honest failures: span errors, poor decomposition, misleading passages.

A language model's knowledge stops at its training cutoff and is patchy on anything long-tail, so a question like "Curious is a women's fragrance by a singer born in what city and state?" is a trap. It needs two facts chained together, and the model is unlikely to hold both reliably. The obvious fix is to give it a search engine. How it should use that engine is the hard part.

Before Search-R1 there were two families of answer. Retrieval-augmented generation (RAG) runs one retrieval using the user's question as the query and pastes the passages into context. That is cheap and robust, but the query is fixed before the model has thought about anything, so it suits single-hop questions and can fall apart on multi-hop ones, where the second query depends on what the first one returned. Search-as-a-tool methods let the model issue its own queries mid-reasoning. IRCoT and ReAct do this with prompting, which the paper calls suboptimal because the LLM was never optimised to interact with a search engine. Toolformer does it with supervised fine-tuning, which needs large sets of high-quality annotated search trajectories that are expensive to produce at scale. And search is non-differentiable, so you can't simply backpropagate through it.

Reinforcement learning sidesteps both problems. DeepSeek-R1 showed that RL from outcome rewards alone can teach a model to reason, with no annotated chains of thought. Search-R1 asks whether the same recipe can teach a model when and what to search. The authors identify three challenges: (1) framework and stability — how to put a search engine inside an RL loop when part of the trajectory is text the model did not write; (2) multi-turn interleaving — the model should decide dynamically how many searches a question needs; and (3) reward design — whether a simple exact-match outcome reward is enough to induce meaningful search behaviour.

The paper answers all three, and describes itself as an extension of DeepSeek-R1-Zero from parametric reasoning to retrieval-driven reasoning.

Key Points

1

RAG retrieves once using the input as the query, so the LLM never chooses what to look up and cannot follow up on what it found

2

Prompted tool use (IRCoT, ReAct) lets the model search mid-reasoning, but nothing optimises how it searches; prompting-based approaches often struggle to generalise

3

SFT tool use (Toolformer) needs large-scale, high-quality annotated trajectories, and the search call itself is non-differentiable

4

RL with outcome rewards, as in DeepSeek-R1, needs neither annotated trajectories nor gradients through the search engine

5

Three open challenges: stable RL with retrieved context, multi-turn interleaved search, and whether a simple outcome reward is sufficient

6

Search-R1 = DeepSeek-R1-Zero-style RL plus a search engine in the rollout loop