PIXELBANKv9.1.0
Menu
Back to Concepts
Large Language Models2025

Qwen3

Unified Thinking and Non-Thinking Modes, Thinking Budgets, and Strong-to-Weak Distillation in an Open Dense + MoE LLM Family

Qwen Team

Read the Paper on arXiv

Paper Overview

Qwen3 (Qwen Team, Alibaba, 2025) is a family of open-weight LLMs — six dense models from 0.6B to 32B and two Mixture-of-Experts models, Qwen3-30B-A3B and the flagship Qwen3-235B-A22B — built around one idea: a single model should be able to both think and not think.

Until then, reasoning and chat were separate products. You ran a chat model for fast answers and a reasoning model (like QwQ-32B) for hard problems, and chose between them per query. Qwen3 folds both into one set of weights, switched per turn with /think and /no_think, and adds a thinking budget that caps reasoning tokens so latency and accuracy can be traded continuously. Remarkably, the budget is not trained — it emerges from teaching the model both endpoints.

The machinery behind this is a 36T-token, 119-language pre-training run in three stages; a four-stage post-training pipeline (long-CoT cold start, reasoning RL, thinking-mode fusion, general RL) for the flagships; and strong-to-weak distillation that passes both modes to small models at about a tenth of the GPU hours.

The flagship reaches 85.7 on AIME'24, 81.5 on AIME'25, 70.7 on LiveCodeBench v5 and a 2,056 CodeForces rating in thinking mode, beats DeepSeek-R1 on 17 of 23 benchmarks with 22B activated parameters, and everything ships under Apache 2.0.

Chapter Roadmap

Click any topic to jump in

1
One Model, Two Modes

Chat and reasoning used to be separate models. Qwen3 folds both into one set of weights, switched per turn by /think and /no_think.

built on

Model and data

2
Dense + MoE Family

Six dense and two MoE models; QK-Norm replaces QKV-bias, and 128 experts with 8 active and no shared expert.

3
36T-Token Pre-training

119 languages, VLM-extracted PDFs and synthetic data; three stages for breadth, reasoning depth, then 32K context.

post-trained into
4
Stages 1–2: Build the Reasoner

A deliberately small long-CoT cold start, then GRPO on 3,995 query-verifier pairs — AIME'24 70.1 → 85.1 in 170 steps.

fused with non-thinking

Two modes, one dial

5
Stage 3: Thinking Mode Fusion

SFT with a chat template where non-thinking answers keep an empty <think></think> block; the model obeys the last flag.

6
The Thinking Budget

Cut the reasoning at a token threshold, insert a stop instruction, and the model answers from partial thinking — untrained, emergent.

broadened and handed down

Flagships and lightweight models

7
Stage 4: General RL

Rewards over 20+ tasks lift mode switching to 98.9 and tool use by +15, at the cost of a few points of peak reasoning.

8
Strong-to-Weak Distillation

Small models learn both modes from the flagships; on-policy distillation beats RL on Qwen3-8B at 1/10 of the GPU hours.

produces

Outcomes and caveats

9
Results

AIME'24 85.7, LiveCodeBench 70.7, CodeForces 2,056; beats DeepSeek-R1 on 17/23 and GPT-4o on 18/23 benchmarks.

10
Long Context & Limits

On RULER, thinking mode slightly hurts retrieval — a reminder that more test-time compute is not always better.

By early 2025 an open-model user who wanted both a fast chat assistant and a careful reasoner had to run two different models. The Qwen line itself was split this way: Qwen2.5 for chat, QwQ-32B for long chain-of-thought reasoning. The report names the same split in the closed world — chat-optimised GPT-4o on one side, dedicated reasoning models on the other. Every application had to decide up front which model a query deserved, and pay to deploy both.

That split is wasteful in an obvious way and a subtle one. The obvious waste is operational: two sets of weights, two serving stacks, a router in front. The subtle waste is that most queries sit somewhere between trivial and olympiad-hard. A greeting does not need thousands of reasoning tokens; a competition problem cannot be solved without them; and a large middle band would benefit from some thinking but not an unbounded amount. A binary choice between models cannot express that middle band at all.

Qwen3's central design decision is to fold both behaviours into a single set of weights. The same model answers immediately when asked with /no_think and reasons at length when asked with /think (or by default). Because one model has learned both endpoints, it also handles the in-between: stop its reasoning early at a user-chosen thinking budget, and it writes the best answer it can from the partial reasoning. Compute per query becomes a dial rather than a model choice.

The rest of the report is the machinery that makes this possible: a family of eight dense and MoE models pre-trained on 36T tokens in 119 languages, a four-stage post-training pipeline that first builds a reasoner and then teaches it to not reason on request, and a strong-to-weak distillation path that hands both modes down to small models at a fraction of the cost. The flagship Qwen3-235B-A22B — 235B parameters, 22B active per token — reaches 85.7 on AIME'24, 81.5 on AIME'25 and 70.7 on LiveCodeBench v5 in thinking mode, and everything is released under Apache 2.0.

Key Points

1

Before Qwen3, reasoning and chat were separate models (e.g. Qwen2.5 vs QwQ-32B; GPT-4o vs dedicated reasoning models) — users switched models to switch behaviour

2

Qwen3 integrates thinking mode (multi-step reasoning) and non-thinking mode (rapid responses) into one set of weights

3

Mode is selected per turn with /think and /no_think flags in the user query or system message; thinking is the default

4

A thinking budget lets users cap reasoning tokens, trading latency for accuracy continuously instead of choosing between two models

5

Family: 6 dense models (0.6B–32B) and 2 MoE models (30B-A3B, 235B-A22B), all open-weight under Apache 2.0

6

Multilingual coverage expands from 29 to 119 languages and dialects compared to Qwen2.5