Skip to content
ContentLora

    Tip: press / anywhere to search.

    concept

    Test-time compute

    Also known as inference-time compute, inference-time scaling, test-time scaling

    Test-time compute, also called inference-time compute, is the computation a model uses when it answers rather than when it is trained. Reasoning models use more of it to generate intermediate steps before a final answer,[1] and a 2024 study found that spending it well can rival using a much larger model.[2]

    Editor reviewedUpdated Frontier AIArtificial intelligenceComputing
    Key facts

    Test-time compute is a scaling lever alongside training scale. For years, progress came mainly from training bigger models on more data with more compute.[3] Labs now also improve results by letting a model compute for longer at answer time, generating intermediate steps before it responds.[1]

    How it works

    A model can spend extra inference compute in two broad ways. It can write a longer chain of thought, a sequence of intermediate reasoning steps that a 2022 study showed improves complex reasoning.[4] Or it can explore several candidate solutions and choose among them: Google DeepMind describes its Deep Think mode as exploring multiple solution paths in parallel.[5] Reasoning models such as openai-o1 and deepseek-r1 are trained with reinforcement learning to make these long chains of reasoning productive.[6][7] The DeepSeek-R1 paper reports that its model learned on its own to spend more thinking time, producing longer responses, as reinforcement learning progressed.[8]

    The evidence

    A 2024 study by Snell and colleagues tested two mechanisms, searching against verifier models and revising answers step by step.[2] Allocating compute adaptively per prompt improved efficiency by more than 4x over a best-of-N baseline, and in FLOPs-matched comparisons a smaller model could beat one 14 times larger on some problems.[2] The gains depended on prompt difficulty, so the best strategy varies from question to question.[9]

    The clearest public example came from ARC-AGI. In December 2024 a preview of OpenAI‘s o3 scored 75.7% on the ARC-AGI-1 semi-private set within the benchmark’s compute limit, and 87.5% in a configuration that used about 172 times more compute.[10]

    Scaling the training behind it

    Using test-time compute well depends on reinforcement learning during training. Epoch AI estimated that OpenAI’s o3 used about ten times the training compute of o1, reached in roughly four months.[11] It projected in May 2025 that this pace would soon run into the limits of total training compute and slow to the overall growth rate of about 4x per year.[12]

    Costs and trade-offs

    More thinking costs money and time. The ARC Prize Foundation reported that GPT-6 Astra’s 62.7% run on ARC-AGI-3 with its standard harness cost $26,098.[13] More compute is also not a cure-all: Apple researchers found that reasoning models’ effort rises with problem complexity up to a point and then declines.[14] For a wider look at these limits, see the analysis of reasoning models’ limits, linked below.[15]

    What it means for evaluation

    If accuracy depends on how long a model may think, a single benchmark number can mislead: the ARC Prize Foundation argued after o1’s release that no score is objective when it is relative to the compute allowed.[16] In 2026 the UK AI Security Institute found the same problem in agent testing.[17] In March it reported, with Irregular, that standard budgets may underestimate models’ ceiling on cyber tasks.[18] In July it concluded more broadly that fixed-budget evaluations can systematically underestimate frontier agents, especially newer ones.[17] The compute a task needs rises with how long it would take a skilled human, so the longest and hardest tasks are the first to be cut off.[19] That matters for measures such as METR‘s time horizons, which rank models by the length of task they can finish.[20]

    Test-time compute in 2026 products

    Labs now ship explicit reasoning modes. Google reported in November 2025 that Gemini 3 Pro scored 37.5% on Humanity’s Last Exam without tools,[21] while its Deep Think mode, which spends more reasoning effort, scored 41.0% and reached 45.1% on ARC-AGI-2.[22] DeepSeek’s V4 models offer separate thinking and non-thinking modes,[23] and OpenAI’s GPT-6 Astra reasons through an extended chain of thought before answering.[24]

    Questions readers ask

    What is test-time compute?

    The computing power a model spends while answering. Reasoning models use more of it to generate intermediate steps before giving a final answer.[1]

    Can extra thinking time replace a bigger model?

    Sometimes. A 2024 study found that a smaller model given well-allocated inference compute could outperform a model 14 times larger on some problems.[2]

    Does more thinking time always help?

    No. The benefit depends on how hard the prompt is, and one study found reasoning effort declines past a certain problem complexity.[9][14]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24

    2. [2]

      A 2024 study of two mechanisms, search against process-based verifier models and adaptive revision of responses, found that allocating test-time compute adaptively per prompt improved test-time compute efficiency by more than 4x over a best-of-N baseline, and that in some settings a smaller model given extra inference compute could outperform a 14x larger model. confirmedas of 2024-08-06

    3. [3]

      A 2020 study found that language-model loss falls as a power law as model size, dataset size and training compute grow, with some trends spanning more than seven orders of magnitude. confirmedas of 2020-01-23

    4. [4]

      A 2022 study showed that prompting large language models to generate a chain of thought, a series of intermediate reasoning steps, significantly improves their performance on complex reasoning. confirmedas of 2022-01-28

    5. [5]

      Google DeepMind described Deep Think as exploring multiple solution paths in parallel and said the IMO model was trained with reinforcement learning techniques focused on multi-step problem solving. confirmedas of 2025-07-21

    6. [6]

      OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21

    7. [7]

      DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04

    8. [8]

      The DeepSeek-R1 paper reports that R1-Zero learned to spend more thinking time, producing longer responses, as reinforcement learning progressed. confirmedas of 2025-09-17

    9. [9]

      The same 2024 study found that how much extra inference compute helps depends on the difficulty of the prompt. confirmedas of 2024-08-06

    10. [10]

      In December 2024 the ARC Prize Foundation reported that a preview of OpenAI's o3 scored 75.7% on the ARC-AGI-1 semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute. confirmedas of 2024-12-20

    11. [11]

      Epoch AI estimated in May 2025 that OpenAI's o3 represented about a 10x scale-up in training compute over o1, reached in roughly four months. reportedas of 2025-05-09

    12. [12]

      Epoch AI projected in May 2025 that if reasoning training kept scaling 10x every few months it would reach the frontier of total training compute within about a year, after which its growth would slow to the overall rate of roughly 4x per year. reportedas of 2025-05-09· forecast

    13. [13]

      On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03

    14. [14]

      The same study found that reasoning models' reasoning effort rises with problem complexity up to a point and then declines. confirmedas of 2025-11-20

    15. [15]

      A 2025 Apple study using controllable puzzles found that large reasoning models face a complete accuracy collapse beyond certain problem complexities. confirmedas of 2025-11-20

    16. [16]

      The ARC Prize Foundation argued after o1's release that when AI systems may use a variable amount of test-time compute, no single benchmark score is objective, because accuracy depends on the compute allowed. confirmedas of 2024-09-13

    17. [17]

      AISI reported in July 2026 that fixed-budget evaluations can systematically underestimate frontier agents' capabilities, because the test-time compute an agent may spend is a major driver of its measured capability. confirmedas of 2026-07-02

    18. [18]

      AISI and Irregular reported in March 2026 that standard evaluation budgets may underestimate models' ceiling on cyber tasks, and that since November 2025 extra budget has increasingly changed results. confirmedas of 2026-03-05

    19. [19]

      AISI found that the compute a task demands rises with how long it would take a skilled human, so the longest and hardest tasks are the first to be cut off by fixed budgets. confirmedas of 2026-07-02

    20. [20]

      METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18

    21. [21]

      Google launched Gemini 3 on 18 November 2025, reporting 37.5% on Humanity's Last Exam without tools and 91.9% on GPQA Diamond for Gemini 3 Pro. confirmedas of 2025-11-18

    22. [22]

      Google reported in November 2025 that Gemini 3 Deep Think scored 45.1% on ARC-AGI-2 and 41.0% on Humanity's Last Exam. confirmedas of 2025-11-18

    23. [23]

      DeepSeek V4 offers both a thinking and a non-thinking mode, and its weights were released openly on Hugging Face. confirmedas of 2026-04-24

    24. [24]

      OpenAI's system card for GPT-6 Astra, dated 3 September 2026, calls it the most capable model OpenAI has ever broadly deployed and says it reasons through an extended chain of thought trained with reinforcement learning. confirmedas of 2026-09-03

    Revision history (2)
    1. Page created.
    2. Added o3's compute-dependent ARC-AGI results, DeepSeek-R1's growing thinking time, and the UK AI Security Institute's 2026 findings that fixed compute budgets understate agents' capabilities.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Test-time compute." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/test-time-compute

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.