PIXELBANKv9.1.0
Menu
Paper breakdownLarge Language ModelsFREE2024

Qwen2.5

18T-Token Pre-training, Hyper-parameter Scaling Laws, and SFT + DPO + GRPO Post-training for an Open 0.5B–72B LLM Family

Qwen Team

10

Concepts

12

Formulas

8

Figures

5

Demos

Paper Overview

Qwen2.5 (Qwen Team, Alibaba — arXiv 2412.15115, Dec 2024) has an unusual thesis for a technical report: the architecture is nearly the same as last time, and that is intentional. Qwen2.5 keeps Qwen2's dense decoder — GQA, SwiGLU, RoPE, QKV bias, RMSNorm — and puts its effort into the parts of the pipeline other than the architecture.

On the pre-training side, the dataset grows from 7T to 18T tokens, built by previous-generation Qwen models that score, synthesise and rebalance web data. Scaling laws are used to predict optimal learning rate and batch size, not only model size, across nine production models. Context is extended in stages, with RoPE base changes and training-free YaRN + Dual Chunk Attention at inference.

On the post-training side, over 1M SFT examples are filtered by verifiers, then RL runs in two stages. Offline DPO handles domains where answers can be checked mechanically, and online GRPO with a learned reward model handles helpfulness, truthfulness and harmlessness.

The open release spans 0.5B to 72B (base and instruct, plus quantized variants), alongside two API-only MoE models, Qwen2.5-Turbo (1M-token context) and Qwen2.5-Plus. The flagship Qwen2.5-72B-Instruct beats Llama-3.1-405B-Instruct on MATH (83.1 vs 73.8), LiveCodeBench (55.5 vs 41.6) and Arena-Hard (81.2 vs 69.3) at about a fifth of its size. Qwen2.5 also became the base for Qwen2.5-Math, Qwen2.5-Coder, QwQ and Qwen2.5-VL.

Explainer video

Chapter Roadmap

Click any topic to jump in

Interactive demo
1
7T → 18T Tokens

Qwen2-scored filtering, Math/Coder corpora, reward-filtered synthetic data and domain rebalancing make 18T tokens worth training on.

trains

Architecture and tuning

2
Model Family

Seven dense sizes from 0.5B to 72B with GQA, SwiGLU, RoPE and QKV bias, plus API-only MoE models Turbo and Plus.

3
Hyper-parameter Scaling Laws

Optimal learning rate and batch size predicted from model size and data size, fitted on 44M–14B runs.

extended by
4
Long-Context Pre-training

4K → 32K with RoPE base 10^6; Turbo climbs to 262K; YaRN + DCA add 4x at inference.

post-trained via
5
Supervised Fine-tuning

Over 1M verified examples across nine skill areas, with outputs up to 8K tokens.

refined by

Two-stage RL

6
Offline RL (DPO)

~150K pass/fail pairs graded by execution and answer matching, where reward models struggle.

7
Online RL (GRPO)

A six-criterion reward model, 8 samples per query, and high-variance queries trained first.

scrutinised and stretched

Signals and scale

8
Reward Model Evaluation

No RM wins every benchmark, and RM scores fail to predict the quality of the RL policy.

9
1M-Token Context

Two-stage long SFT, short-only RL, 100% passkey at 1M, and 3.2–4.3x faster TTFT via sparse attention.

measured in
10
Results

72B-Instruct beats Llama-3.1-405B-Instruct on 7 of 13 instruct benchmarks at about 1/5.6 of its size.

Concept by concept

10 sections, in reading order.

01 / 10Better Data — From 7T to 18T Tokens

Qwen2.5 keeps Qwen2's architecture almost untouched. What changes is the data, and the report is blunt about where it thinks capability comes from: Figure 1 plots the Qwen series against pre-training tokens — 3T for Qwen1.5, 7T for Qwen2, 18T for Qwen2.5 — and credits "scale together with mixture" for the gains.

Scale alone would be easy to describe and hard to make useful. Web-scale crawls are dominated by whatever the internet produces most of, and the paper names it: e-commerce, social media and entertainment are heavily over-represented, often as repetitive, template-based or machine-generated text, while technology, science and academic research — the domains that carry dense information — are under-represented. Simply collecting 2.5x more of the same distribution would mostly buy 2.5x more product listings.

So the 18T tokens are the output of four interventions. Filtering: Qwen2-Instruct models score every sample along multiple dimensions, which works better than Qwen2's filter because Qwen2 itself was trained on a larger multilingual corpus. Domain data: the math and code corpora built for Qwen2.5-Math and Qwen2.5-Coder are folded directly into general pre-training. Synthetic data: Qwen2-72B-Instruct and Qwen2-Math-72B-Instruct generate math, code and knowledge text, which is then filtered by a general reward model and by Qwen2-Math-RM-72B. Mixture: Qwen2-Instruct classifies content by domain, over-represented domains are down-sampled and high-value domains up-sampled.

The pattern worth noticing is self-bootstrapping. Every one of the four steps uses a previous-generation Qwen model as a judge, a generator or a classifier. The model family is now a component of its own data pipeline — the same move that lets later models in the series (QwQ, Qwen2.5-Math, Qwen2.5-VL) build on Qwen2.5 in turn.

Better Data — From 7T to 18T Tokens diagram

Key Points

  • 1

    Pre-training data grows from 7 trillion tokens (Qwen2) to 18 trillion tokens (Qwen2.5), with a focus on knowledge, coding and mathematics

  • 2

    Better filtering: Qwen2-Instruct models act as multi-dimensional quality scorers — improving both retention of good data and removal of bad data across languages

  • 3

    Better math and code: training data from Qwen2.5-Math and Qwen2.5-Coder is integrated directly into general pre-training

  • 4

    Better synthetic data: generated by Qwen2-72B-Instruct and Qwen2-Math-72B-Instruct, then filtered by a proprietary general reward model and Qwen2-Math-RM-72B

  • 5

    Better mixture: e-commerce, social media and entertainment are down-sampled; technology, science and academic research are up-sampled

  • 6

    On base models the effect is visible in Table 2: Qwen2-72B → Qwen2.5-72B moves MMLU 84.2 → 86.1, MATH 50.9 → 62.1, MBPP 76.9 → 84.7