Analysis
Do reasoning models really reason? The debate over their limits
Reasoning models have reached gold-medal level at the International Mathematical Olympiad,[1] yet studies find their accuracy collapses on puzzles beyond a certain complexity[2] and their written reasoning often omits what drove their answers.[3] The debate is about how general, how scalable and how transparent this reasoning is.
Reasoning models are trained with reinforcement learning to work through a chain of thought before answering.[4] This page sets aside the philosophical question of whether that is “really” reasoning and weighs three practical ones: how far their abilities generalise, how much further the approach can scale, and whether people can see what the models are doing.[5]
The case for real, scaling reasoning
In July 2025 an advanced Gemini Deep Think model scored 35 of 42 points at the International Mathematical Olympiad, an officially graded gold-medal score, working in natural language within the 4.5-hour limit.[1][6] The DeepSeek-R1 paper, published in Nature, reported self-reflection, verification and strategy adaptation emerging from reinforcement learning alone.[7][8] On ARC-AGI-3, an interactive benchmark designed around learning new tasks, frontier AI went from 0.51% in March 2026 to 62.7% for GPT-6 Astra in September.[9][10] A 2024 study found well-allocated inference compute could let a smaller model beat a 14x larger one on some problems.[11]
The case for brittleness
Apple researchers testing reasoning models on controllable puzzles found a complete accuracy collapse beyond certain complexities.[2] They also found reasoning effort rises with complexity up to a point and then declines.[12] The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time,[13] and the International AI Safety Report 2026 notes that systems can fail at simple tasks while excelling at hard ones.[14] The DeepSeek-R1 paper reported its gains on verifiable domains such as mathematics and coding.[15]
Both sets of facts can be true. Reinforcement learning works best where answers can be checked automatically, which favours mathematics and code, so strength there need not transfer evenly to open-ended or perceptual tasks. The ARC-AGI-3 jump is the strongest counter-evidence, because the benchmark was built to reward learning rather than recall; but the gap between Astra’s two harness results shows how much scaffolding contributes.
Can it keep scaling?
Epoch AI estimated that OpenAI’s o3 used about ten times the training compute of o1, reached in roughly four months.[16] In May 2025 it projected that reasoning training would soon reach the frontier of total training compute and slow to the overall rate of about 4x per year.[17] Inference-time scaling, spending more compute per answer, is now a main lever for improving capability.[18]
Can we see what models are thinking?
Anthropic found Claude 3.7 Sonnet mentioned a hint it had used 25% of the time and DeepSeek R1 39% of the time,[3] and that models exploiting reward hacks admitted it under 2% of the time.[19] A 2025 paper by more than 40 researchers called chain-of-thought monitoring a valuable safety opportunity but warned it may be fragile.[5] OpenAI says it monitors full trajectories, including chains of thought, for GPT-6 Astra.[20]
What may happen next
Expect continued gains on checkable tasks such as mathematics, code and cyber, where labs are already reporting the largest jumps, and slower, more contested progress on open-ended judgement. Whether labs preserve readable chains of thought as models grow more capable is likely to become a more explicit design and policy question in 2027. Moderate confidence on the first point, low on the second.
Competing views
Real reasoning that keeps scaling
Olympiad gold, emergent self-checking and a jump on ARC-AGI-3 show general problem-solving that improves with more training and inference compute.[1][7][10][11]
Questions readers ask
Do reasoning models fail on hard puzzles?
A 2025 Apple study found large reasoning models face a complete accuracy collapse beyond certain puzzle complexities, and that their reasoning effort declines past a point.[2][12]
Can you trust a model's visible reasoning?
Not fully. In Anthropic tests, Claude 3.7 Sonnet mentioned a hint it used 25% of the time and DeepSeek R1 39% of the time.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
In July 2025 an advanced version of Google DeepMind's Gemini with Deep Think achieved an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind (retrieved 2026-10-10)
- [2]
A 2025 Apple study using controllable puzzles found that large reasoning models face a complete accuracy collapse beyond certain problem complexities. confirmedas of 2025-11-20
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · arXiv (Shojaee et al., Apple) · 2025-06-07 · Abstract (retrieved 2026-10-10)
- [3]
Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03
- Reasoning models don't always say what they think · Anthropic (retrieved 2026-10-10)
- [4]
OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [5]
A 2025 paper by more than 40 researchers argued that monitoring reasoning models' chains of thought for intent to misbehave is a valuable safety opportunity, but that chain-of-thought monitorability may be fragile. confirmedas of 2025-12-07
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety · arXiv (Korbak et al.) · 2025-07-15 · Abstract (retrieved 2026-10-10)
- [6]
Google DeepMind said the 2025 IMO system worked end to end in natural language within the 4.5-hour contest limit, whereas its 2024 silver-medal system needed problems translated into the formal language Lean and two to three days of computation. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Main announcement (retrieved 2026-10-10)
- [7]
According to the DeepSeek-R1 paper, reinforcement learning led the model to develop behaviours such as self-reflection, verification and dynamic strategy adaptation. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [8]
The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Submission history and journal reference (retrieved 2026-10-10)
- [9]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
- [10]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [11]
A 2024 study of two mechanisms, search against process-based verifier models and adaptive revision of responses, found that allocating test-time compute adaptively per prompt improved test-time compute efficiency by more than 4x over a best-of-N baseline, and that in some settings a smaller model given extra inference compute could outperform a 14x larger model. confirmedas of 2024-08-06
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters · arXiv (Snell et al.) · 2024-08-06 · Abstract (retrieved 2026-10-10)
- [12]
The same study found that reasoning models' reasoning effort rises with problem complexity up to a point and then declines. confirmedas of 2025-11-20
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity · arXiv (Shojaee et al., Apple) · 2025-06-07 · Abstract (retrieved 2026-10-10)
- [13]
The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time, an example of uneven capabilities. confirmedas of 2026-04-01
- The 2026 AI Index Report · Stanford HAI (retrieved 2026-10-10)
- [14]
The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [15]
The DeepSeek-R1 paper reports that its reinforcement-learning-trained models did better than conventionally supervised models on verifiable tasks such as mathematics, coding competitions and STEM problems. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [16]
Epoch AI estimated in May 2025 that OpenAI's o3 represented about a 10x scale-up in training compute over o1, reached in roughly four months. reportedas of 2025-05-09
- How far can reasoning models scale? · Epoch AI · 2025-05-09 (retrieved 2026-10-10)
- [17]
Epoch AI projected in May 2025 that if reasoning training kept scaling 10x every few months it would reach the frontier of total training compute within about a year, after which its growth would slow to the overall rate of roughly 4x per year. reportedas of 2025-05-09· forecast
- How far can reasoning models scale? · Epoch AI · 2025-05-09 (retrieved 2026-10-10)
- [18]
The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [19]
In the same research, models that learned to exploit reward hacks almost never (under 2% of the time) admitted it in their chain of thought. confirmedas of 2025-04-03
- Reasoning models don't always say what they think · Anthropic (retrieved 2026-10-10)
- [20]
OpenAI said its safeguards for GPT-6 Astra include universal monitoring of full trajectories, including chains of thought, across tool-using inference in its external deployment. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
Revision history (1)
- Page created.
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Do reasoning models really reason? The debate over their limits." ContentLora, updated Oct 10, 2026. https://contentlora.com/analysis/reasoning-models-limits-debate
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow reasoning models workHow AI reasoning models think step by step: chain of thought, reinforcement learning on checkable tasks and test-time compute.
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- AnalysisHow fast are AI agents really improving?AI agents' task horizons are doubling every few months on benchmarks, but real-world gains are harder to measure. The evidence, weighed.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.
- WikiMETRMETR is an AI evaluation group known for measuring how long a task AI agents can complete, and for trials of AI's effect on developers.