Explainer
How reasoning models work
Reasoning models are language models trained to work through a problem in a long chain of intermediate steps before answering. Labs train this behaviour with large-scale reinforcement learning,[1][2] and the models get better results by spending more computing power at answer time.[3]
Reasoning models are a central development in frontier AI since 2024. Instead of answering at once, they generate long internal reasoning and are trained, with reinforcement learning, to make that reasoning lead to correct answers.[1] The International AI Safety Report 2026 calls this inference-time scaling, and it has become one of the main ways labs improve capabilities.[3]
Thinking out loud
In 2022, researchers found that if you show a language model worked examples that spell out each step, it starts writing out its own steps and gets much better at multi-step problems.[4] This written working is called a “chain of thought”.[4]
Chain-of-thought prompting supplies a few exemplars containing intermediate reasoning steps; generating such a chain significantly improves performance on complex reasoning tasks.[4] Early on this was a prompting trick applied to a fixed model. Reasoning models make it a trained behaviour.[1]
Training reasoning with reinforcement learning
A reasoning model practises on problems whose answers can be checked, such as maths problems or code that must pass tests, and is rewarded when it gets them right. Over many rounds it learns ways of thinking that work. OpenAI’s openai-o1 series is trained this way,[1] and DeepSeek reported that this kind of training could work even without people writing example reasoning for the model to copy.[2]
The o1 system card states that the series is trained with large-scale reinforcement learning to reason using chain of thought.[1] The deepseek-r1 paper, later published in Nature, reported that pure RL without human-labelled reasoning trajectories elicited reasoning,[2][5] with self-reflection, verification and dynamic strategy adaptation emerging during training.[6] The paper reports gains in verifiable domains such as mathematics, competitive coding and STEM.[7] By 2026 the approach is widespread: OpenAI’s GPT-6 Astra system card describes a model trained through reinforcement learning to produce a long internal chain of thought.[8]
Thinking longer at answer time
Reasoning models can be given more “thinking time” for harder questions. A 2024 study found that using extra computing power wisely at answer time can let a smaller model beat one 14 times its size in some settings.[9] How much this helps depends on how hard the question is.[10]
Test-time compute adds a scaling axis beyond pre-training. Snell et al. found that allocating inference compute adaptively per prompt improved efficiency by more than 4x over best-of-N, and that on FLOPs-matched comparisons a smaller model could outperform a 14x larger one on some problems.[9] The benefit depends on prompt difficulty.[10] Google DeepMind’s Deep Think explores multiple solution paths in parallel rather than one linear chain.[11]
What reasoning models have achieved
In July 2025 an advanced Gemini Deep Think model earned an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems.[12] It worked in natural language within the 4.5-hour limit, whereas the 2024 system needed formal translation and days of computation.[13] In September 2026 OpenAI’s GPT-6 Astra scored 62.7% on ARC-AGI-3, a benchmark on which frontier AI scored 0.51% at its March 2026 launch.[14][15]
Limits and open questions
Reasoning does not remove unevenness. Apple researchers found that reasoning models’ accuracy collapses beyond certain puzzle complexities.[16] Anthropic found that a model’s written reasoning does not always reveal what actually drove its answer.[17] These questions are covered in the analysis page on the limits of reasoning models.[18]
Questions readers ask
What is a chain of thought?
A series of intermediate reasoning steps a model writes out before its final answer. A 2022 study found that generating one significantly improves performance on complex reasoning.[4]
How are reasoning models trained?
With large-scale reinforcement learning. OpenAI says its o1 series is trained this way to reason using chain of thought, and DeepSeek reported that pure reinforcement learning, without human-labelled reasoning examples, could develop reasoning.[1][2]
Why do reasoning models take longer to answer?
They use more computing power at answer time to generate intermediate steps before giving a final answer, an approach called inference-time scaling.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [2]
DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [3]
The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [4]
A 2022 study showed that prompting large language models to generate a chain of thought, a series of intermediate reasoning steps, significantly improves their performance on complex reasoning. confirmedas of 2022-01-28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models · arXiv (Wei et al.) · 2022-01-28 · Abstract (retrieved 2026-10-10)
- [5]
The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Submission history and journal reference (retrieved 2026-10-10)
- [6]
According to the DeepSeek-R1 paper, reinforcement learning led the model to develop behaviours such as self-reflection, verification and dynamic strategy adaptation. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [7]
The DeepSeek-R1 paper reports that its reinforcement-learning-trained models did better than conventionally supervised models on verifiable tasks such as mathematics, coding competitions and STEM problems. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [8]
OpenAI's system card for GPT-6 Astra, dated 3 September 2026, calls it the most capable model OpenAI has ever broadly deployed and says it reasons through an extended chain of thought trained with reinforcement learning. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [9]
A 2024 study of two mechanisms, search against process-based verifier models and adaptive revision of responses, found that allocating test-time compute adaptively per prompt improved test-time compute efficiency by more than 4x over a best-of-N baseline, and that in some settings a smaller model given extra inference compute could outperform a 14x larger model. confirmedas of 2024-08-06
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters · arXiv (Snell et al.) · 2024-08-06 · Abstract (retrieved 2026-10-10)
- [10]
The same 2024 study found that how much extra inference compute helps depends on the difficulty of the prompt. confirmedas of 2024-08-06
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters · arXiv (Snell et al.) · 2024-08-06 · Abstract (retrieved 2026-10-10)
- [11]
Google DeepMind described Deep Think as exploring multiple solution paths in parallel and said the IMO model was trained with reinforcement learning techniques focused on multi-step problem solving. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Technical approach section (retrieved 2026-10-10)
- [12]
In July 2025 an advanced version of Google DeepMind's Gemini with Deep Think achieved an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind (retrieved 2026-10-10)
- [13]
Google DeepMind said the 2025 IMO system worked end to end in natural language within the 4.5-hour contest limit, whereas its 2024 silver-medal system needed problems translated into the formal language Lean and two to three days of computation. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Main announcement (retrieved 2026-10-10)
- [14]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [15]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [16]
A 2025 Apple study using controllable puzzles found that large reasoning models face a complete accuracy collapse beyond certain problem complexities. confirmedas of 2025-11-20
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · arXiv (Shojaee et al., Apple) · 2025-06-07 · Abstract (retrieved 2026-10-10)
- [17]
Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03
- Reasoning models don't always say what they think · Anthropic (retrieved 2026-10-10)
- [18]
The same study found that reasoning models' reasoning effort rises with problem complexity up to a point and then declines. confirmedas of 2025-11-20
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · arXiv (Shojaee et al., Apple) · 2025-06-07 · Abstract (retrieved 2026-10-10)
Revision history (1)
- Page created.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"How reasoning models work." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-reasoning-models-work
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- ExplainerHow large language models workA plain guide to large language models: the Transformer, scaling laws and human-feedback training, at beginner and expert level.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.