Qwen3
Unified Thinking and Non-Thinking Modes, Thinking Budgets, and Strong-to-Weak Distillation in an Open Dense + MoE LLM Family
Qwen Team
Read the Paper on arXivPaper Overview
Qwen3 (Qwen Team, Alibaba, 2025) is a family of open-weight LLMs — six dense models from 0.6B to 32B and two Mixture-of-Experts models, Qwen3-30B-A3B and the flagship Qwen3-235B-A22B — built around one idea: a single model should be able to both think and not think.
Until then, reasoning and chat were separate products. You ran a chat model for fast answers and a reasoning model (like QwQ-32B) for hard problems, and chose between them per query. Qwen3 folds both into one set of weights, switched per turn with /think and /no_think, and adds a thinking budget that caps reasoning tokens so latency and accuracy can be traded continuously. Remarkably, the budget is not trained — it emerges from teaching the model both endpoints.
The machinery behind this is a 36T-token, 119-language pre-training run in three stages; a four-stage post-training pipeline (long-CoT cold start, reasoning RL, thinking-mode fusion, general RL) for the flagships; and strong-to-weak distillation that passes both modes to small models at about a tenth of the GPU hours.
The flagship reaches 85.7 on AIME'24, 81.5 on AIME'25, 70.7 on LiveCodeBench v5 and a 2,056 CodeForces rating in thinking mode, beats DeepSeek-R1 on 17 of 23 benchmarks with 22B activated parameters, and everything ships under Apache 2.0.
Chapter Roadmap
Click any topic to jump in
One Model, Two Modes
Chat and reasoning used to be separate models. Qwen3 folds both into one set of weights, switched per turn by /think and /no_think.
Model and data
Dense + MoE Family
Six dense and two MoE models; QK-Norm replaces QKV-bias, and 128 experts with 8 active and no shared expert.
36T-Token Pre-training
119 languages, VLM-extracted PDFs and synthetic data; three stages for breadth, reasoning depth, then 32K context.
Stages 1–2: Build the Reasoner
A deliberately small long-CoT cold start, then GRPO on 3,995 query-verifier pairs — AIME'24 70.1 → 85.1 in 170 steps.
Two modes, one dial
Stage 3: Thinking Mode Fusion
SFT with a chat template where non-thinking answers keep an empty <think></think> block; the model obeys the last flag.
The Thinking Budget
Cut the reasoning at a token threshold, insert a stop instruction, and the model answers from partial thinking — untrained, emergent.
Flagships and lightweight models
Stage 4: General RL
Rewards over 20+ tasks lift mode switching to 98.9 and tool use by +15, at the cost of a few points of peak reasoning.
Strong-to-Weak Distillation
Small models learn both modes from the flagships; on-policy distillation beats RL on Qwen3-8B at 1/10 of the GPU hours.
Outcomes and caveats
Results
AIME'24 85.7, LiveCodeBench 70.7, CodeForces 2,056; beats DeepSeek-R1 on 17/23 and GPT-4o on 18/23 benchmarks.
Long Context & Limits
On RULER, thinking mode slightly hurts retrieval — a reminder that more test-time compute is not always better.
By early 2025 an open-model user who wanted both a fast chat assistant and a careful reasoner had to run two different models. The Qwen line itself was split this way: Qwen2.5 for chat, QwQ-32B for long chain-of-thought reasoning. The report names the same split in the closed world — chat-optimised GPT-4o on one side, dedicated reasoning models on the other. Every application had to decide up front which model a query deserved, and pay to deploy both.
That split is wasteful in an obvious way and a subtle one. The obvious waste is operational: two sets of weights, two serving stacks, a router in front. The subtle waste is that most queries sit somewhere between trivial and olympiad-hard. A greeting does not need thousands of reasoning tokens; a competition problem cannot be solved without them; and a large middle band would benefit from some thinking but not an unbounded amount. A binary choice between models cannot express that middle band at all.
Qwen3's central design decision is to fold both behaviours into a single set of weights. The same model answers immediately when asked with /no_think and reasons at length when asked with /think (or by default). Because one model has learned both endpoints, it also handles the in-between: stop its reasoning early at a user-chosen thinking budget, and it writes the best answer it can from the partial reasoning. Compute per query becomes a dial rather than a model choice.
The rest of the report is the machinery that makes this possible: a family of eight dense and MoE models pre-trained on 36T tokens in 119 languages, a four-stage post-training pipeline that first builds a reasoner and then teaches it to not reason on request, and a strong-to-weak distillation path that hands both modes down to small models at a fraction of the cost. The flagship Qwen3-235B-A22B — 235B parameters, 22B active per token — reaches 85.7 on AIME'24, 81.5 on AIME'25 and 70.7 on LiveCodeBench v5 in thinking mode, and everything is released under Apache 2.0.
Key Points
Before Qwen3, reasoning and chat were separate models (e.g. Qwen2.5 vs QwQ-32B; GPT-4o vs dedicated reasoning models) — users switched models to switch behaviour
Qwen3 integrates thinking mode (multi-step reasoning) and non-thinking mode (rapid responses) into one set of weights
Mode is selected per turn with /think and /no_think flags in the user query or system message; thinking is the default
A thinking budget lets users cap reasoning tokens, trading latency for accuracy continuously instead of choosing between two models
Family: 6 dense models (0.6B–32B) and 2 MoE models (30B-A3B, 235B-A22B), all open-weight under Apache 2.0
Multilingual coverage expands from 29 to 119 languages and dialects compared to Qwen2.5