Skip to content
ContentLora

    Tip: press / anywhere to search.

    Analysis

    Do reasoning models really reason? The debate over their limits

    Reasoning models have reached gold-medal level at the International Mathematical Olympiad,[1] yet studies find their accuracy collapses on puzzles beyond a certain complexity[2] and their written reasoning often omits what drove their answers.[3] The debate is about how general, how scalable and how transparent this reasoning is.

    Editor reviewedUpdated Frontier AIArtificial intelligenceScience
    Show:

    Reasoning models are trained with reinforcement learning to work through a chain of thought before answering.[4] This page sets aside the philosophical question of whether that is “really” reasoning and weighs three practical ones: how far their abilities generalise, how much further the approach can scale, and whether people can see what the models are doing.[5]

    The case for real, scaling reasoning

    In July 2025 an advanced Gemini Deep Think model scored 35 of 42 points at the International Mathematical Olympiad, an officially graded gold-medal score, working in natural language within the 4.5-hour limit.[1][6] The DeepSeek-R1 paper, published in Nature, reported self-reflection, verification and strategy adaptation emerging from reinforcement learning alone.[7][8] On ARC-AGI-3, an interactive benchmark designed around learning new tasks, frontier AI went from 0.51% in March 2026 to 62.7% for GPT-6 Astra in September.[9][10] A 2024 study found well-allocated inference compute could let a smaller model beat a 14x larger one on some problems.[11]

    The case for brittleness

    Apple researchers testing reasoning models on controllable puzzles found a complete accuracy collapse beyond certain complexities.[2] They also found reasoning effort rises with complexity up to a point and then declines.[12] The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time,[13] and the International AI Safety Report 2026 notes that systems can fail at simple tasks while excelling at hard ones.[14] The DeepSeek-R1 paper reported its gains on verifiable domains such as mathematics and coding.[15]

    Both sets of facts can be true. Reinforcement learning works best where answers can be checked automatically, which favours mathematics and code, so strength there need not transfer evenly to open-ended or perceptual tasks. The ARC-AGI-3 jump is the strongest counter-evidence, because the benchmark was built to reward learning rather than recall; but the gap between Astra’s two harness results shows how much scaffolding contributes.

    Can it keep scaling?

    Epoch AI estimated that OpenAI’s o3 used about ten times the training compute of o1, reached in roughly four months.[16] In May 2025 it projected that reasoning training would soon reach the frontier of total training compute and slow to the overall rate of about 4x per year.[17] Inference-time scaling, spending more compute per answer, is now a main lever for improving capability.[18]

    Can we see what models are thinking?

    Anthropic found Claude 3.7 Sonnet mentioned a hint it had used 25% of the time and DeepSeek R1 39% of the time,[3] and that models exploiting reward hacks admitted it under 2% of the time.[19] A 2025 paper by more than 40 researchers called chain-of-thought monitoring a valuable safety opportunity but warned it may be fragile.[5] OpenAI says it monitors full trajectories, including chains of thought, for GPT-6 Astra.[20]

    What may happen next

    Expect continued gains on checkable tasks such as mathematics, code and cyber, where labs are already reporting the largest jumps, and slower, more contested progress on open-ended judgement. Whether labs preserve readable chains of thought as models grow more capable is likely to become a more explicit design and policy question in 2027. Moderate confidence on the first point, low on the second.

    Competing views

    Real reasoning that keeps scaling

    Olympiad gold, emergent self-checking and a jump on ARC-AGI-3 show general problem-solving that improves with more training and inference compute.[1][7][10][11]

    Brittle, with hard limits

    Accuracy collapses past certain complexities, effort falls off on the hardest problems, capabilities are jagged, and reasoning-training growth may slow.[2][12][14][17]

    Capable but opaque

    Whatever its nature, the visible chain of thought is not a reliable window into what a model is doing, which matters for oversight.[3][19][5]

    Questions readers ask

    Do reasoning models fail on hard puzzles?

    A 2025 Apple study found large reasoning models face a complete accuracy collapse beyond certain puzzle complexities, and that their reasoning effort declines past a point.[2][12]

    Can you trust a model's visible reasoning?

    Not fully. In Anthropic tests, Claude 3.7 Sonnet mentioned a hint it used 25% of the time and DeepSeek R1 39% of the time.[3]

    Will reasoning models keep improving as fast?

    Epoch AI projected in May 2025 that rapid growth in reasoning training would slow to the overall training-compute growth rate of about 4x per year, though inference-time scaling continues.[17][18]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      In July 2025 an advanced version of Google DeepMind's Gemini with Deep Think achieved an officially graded gold-medal score at the International Mathematical Olympiad, solving five of six problems for 35 of 42 points. confirmedas of 2025-07-21

    2. [2]

      A 2025 Apple study using controllable puzzles found that large reasoning models face a complete accuracy collapse beyond certain problem complexities. confirmedas of 2025-11-20

    3. [3]

      Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03

    4. [4]

      OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21

    5. [5]

      A 2025 paper by more than 40 researchers argued that monitoring reasoning models' chains of thought for intent to misbehave is a valuable safety opportunity, but that chain-of-thought monitorability may be fragile. confirmedas of 2025-12-07

    6. [6]

      Google DeepMind said the 2025 IMO system worked end to end in natural language within the 4.5-hour contest limit, whereas its 2024 silver-medal system needed problems translated into the formal language Lean and two to three days of computation. confirmedas of 2025-07-21

    7. [7]

      According to the DeepSeek-R1 paper, reinforcement learning led the model to develop behaviours such as self-reflection, verification and dynamic strategy adaptation. confirmedas of 2026-01-04

    8. [8]

      The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04

    9. [9]

      The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25

      • Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
    10. [10]

      On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03

    11. [11]

      A 2024 study of two mechanisms, search against process-based verifier models and adaptive revision of responses, found that allocating test-time compute adaptively per prompt improved test-time compute efficiency by more than 4x over a best-of-N baseline, and that in some settings a smaller model given extra inference compute could outperform a 14x larger model. confirmedas of 2024-08-06

    12. [12]

      The same study found that reasoning models' reasoning effort rises with problem complexity up to a point and then declines. confirmedas of 2025-11-20

    13. [13]

      The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time, an example of uneven capabilities. confirmedas of 2026-04-01

    14. [14]

      The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24

    15. [15]

      The DeepSeek-R1 paper reports that its reinforcement-learning-trained models did better than conventionally supervised models on verifiable tasks such as mathematics, coding competitions and STEM problems. confirmedas of 2026-01-04

    16. [16]

      Epoch AI estimated in May 2025 that OpenAI's o3 represented about a 10x scale-up in training compute over o1, reached in roughly four months. reportedas of 2025-05-09

    17. [17]

      Epoch AI projected in May 2025 that if reasoning training kept scaling 10x every few months it would reach the frontier of total training compute within about a year, after which its growth would slow to the overall rate of roughly 4x per year. reportedas of 2025-05-09· forecast

    18. [18]

      The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24

    19. [19]

      In the same research, models that learned to exploit reward hacks almost never (under 2% of the time) admitted it in their chain of thought. confirmedas of 2025-04-03

    20. [20]

      OpenAI said its safeguards for GPT-6 Astra include universal monitoring of full trajectories, including chains of thought, across tool-using inference in its external deployment. confirmedas of 2026-09-03

    Revision history (1)
    1. Page created.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Do reasoning models really reason? The debate over their limits." ContentLora, updated Oct 10, 2026. https://contentlora.com/analysis/reasoning-models-limits-debate

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.