product
OpenAI o1
Also known as o1, o1-preview
OpenAI o1 is a series of language models trained with large-scale reinforcement learning to reason using a chain of thought before they answer.[1] OpenAI's later o3 and GPT-6 models also rely on reinforcement-learning-trained reasoning.[2][3]
Key facts
OpenAI o1 is the reference point for reasoning models in this crash course. Its system card describes a model series trained with large-scale reinforcement learning to reason using chain of thought.[1] Instead of producing an answer straight away, o1 works through a long internal sequence of steps, using more compute at answer time.[4]
Release and testing
OpenAI released o1-preview and o1-mini in September 2024; the ARC Prize Foundation, writing on 13 September, said it had gained access to the newly released models over the previous 24 hours.[5] The full o1 followed on 5 December 2024.[6] Before that launch, the UK and US AI safety institutes jointly tested o1’s cyber, biological, and software and AI development capabilities and shared their findings with OpenAI.[7] The system card covers OpenAI’s own safety evaluations, external red teaming and Preparedness Framework evaluations for o1 and o1-mini.[8]
What o1 introduced
Chain-of-thought prompting, asking a model to write out intermediate steps, had been known since 2022.[9] o1 turned this into a trained behaviour: reinforcement learning rewards reasoning that leads to correct answers.[1] This made test-time-compute a lever labs could pull, spending more computation per answer to get better results.[4]
Safety implications
OpenAI’s system card argues that reasoning helps safety, because the model can reason about safety policies when it responds to a request.[10] OpenAI calls this approach deliberative alignment.[11] It also says that chain-of-thought training introduces new risks that need robust alignment methods.[10] Later research found that written reasoning does not always reflect what drives a model’s answer, a debate covered in the analysis of reasoning models’ limits.[12]
Successors
Epoch AI estimated that o3 used about ten times the training compute of o1, reached in roughly four months.[2] In December 2024 a preview of o3 scored 75.7% on ARC-AGI-1 within the benchmark’s compute limit and 87.5% with about 172 times more compute.[13] The ARC Prize Foundation had already argued, after o1’s release, that a single score means little when accuracy depends on how much test-time compute a model is allowed.[14] Under its January 2026 methodology, METR estimated a 50% time horizon of about 121 minutes for o3, against about 214 minutes for GPT-5.[15] By September 2026 OpenAI’s flagship was GPT-6 Astra, which its system card says reasons with an extended chain of thought trained with reinforcement learning, and which OpenAI rated at a Critical level of cybersecurity capability.[3][16]
Competitors
About a month after o1’s system card, DeepSeek posted deepseek-r1, which reported that pure reinforcement learning could develop reasoning without human-labelled reasoning examples.[17][18] Google’s Gemini Deep Think took a parallel approach, exploring multiple solution paths at once.[19]
OpenAI’s reasoning models in 2026
OpenAI’s GPT-5.6 family of Sol, Terra and Luna, released in July 2026, was treated as High capability for cybersecurity and for biological and chemical risk.[20] OpenAI said Sol and Terra could find vulnerabilities and pieces of exploits but could not carry out end-to-end attacks on hardened targets in its tests.[21] Its deployment safety hub lists system cards for GPT-5.6 (9 July 2026), GPT-6 Astra (3 September 2026), a GPT-6.1 Sol addendum (29 September 2026) and GPT-6 Sol and Luna (7 October 2026).[22] OpenAI rated Sol and Luna High, but below Critical, for cybersecurity and biological and chemical capability,[23] and said neither reaches its High threshold for AI self-improvement.[24] GPT-6.1 Sol, by contrast, is treated as Critical for cybersecurity, like Astra.[25] For Astra, OpenAI says it monitors full trajectories, including chains of thought, across tool-using inference in its external deployment.[26] On ARC-AGI-3, Astra scored 62.7% with a standard harness, against 0.51% for frontier AI at the benchmark’s March 2026 launch.[27][28]
Questions readers ask
What made o1 different from earlier chatbots?
It was trained with large-scale reinforcement learning to reason through a chain of thought before answering, rather than replying immediately.[1]
Is o1 still OpenAI's main model?
No. By September 2026 OpenAI had released GPT-6 Astra, which its system card calls the most capable model OpenAI has ever broadly deployed, and which also reasons with an extended chain of thought.[3]
How much bigger was o3 than o1?
Epoch AI estimated that o3 used about ten times the training compute of o1, reached in roughly four months.[2]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [2]
Epoch AI estimated in May 2025 that OpenAI's o3 represented about a 10x scale-up in training compute over o1, reached in roughly four months. reportedas of 2025-05-09
- How far can reasoning models scale? · Epoch AI · 2025-05-09 (retrieved 2026-10-10)
- [3]
OpenAI's system card for GPT-6 Astra, dated 3 September 2026, calls it the most capable model OpenAI has ever broadly deployed and says it reasons through an extended chain of thought trained with reinforcement learning. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [4]
The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24
- International AI Safety Report 2026 · International AI Safety Report (UK DSIT 2026/001) · 2026-02-24 (retrieved 2026-10-10)
- [5]
OpenAI released its first reasoning models, o1-preview and o1-mini, on 12 September 2024; the ARC Prize Foundation, writing on 13 September, said it had gained access to the newly released models over the previous 24 hours. confirmedas of 2024-09-13
- OpenAI o1 Results on ARC-AGI-Pub · ARC Prize Foundation · 2024-09-13 · Post dated 13 Sep 2024 (retrieved 2026-10-10)
- [6]
The full version of OpenAI o1 was released on 5 December 2024. confirmedas of 2024-12-05
- OpenAI o1 Results on ARC-AGI-Pub · ARC Prize Foundation · 2024-09-13 (retrieved 2026-10-10)
- [7]
The UK and US AI Safety Institutes jointly evaluated OpenAI's o1 before its December 2024 release, testing cyber, biological, and software and AI development capabilities and sharing findings with OpenAI before launch. confirmedas of 2024-12-18
- Pre-Deployment evaluation of OpenAI's o1 model · UK AI Security Institute · 2024-12-18 (retrieved 2026-10-10)
- [8]
The o1 system card covers safety evaluations, external red teaming and Preparedness Framework evaluations for OpenAI o1 and o1-mini. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 (retrieved 2026-10-10)
- [9]
A 2022 study showed that prompting large language models to generate a chain of thought, a series of intermediate reasoning steps, significantly improves their performance on complex reasoning. confirmedas of 2022-01-28
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models · arXiv (Wei et al.) · 2022-01-28 · Abstract (retrieved 2026-10-10)
- [10]
The o1 system card says the models can reason about safety policies when responding to requests, and that chain-of-thought training brings both safety benefits and new risks that need robust alignment methods. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 · Abstract (retrieved 2026-10-10)
- [11]
OpenAI says o1 models can reason about its safety policies in context when answering potentially unsafe prompts, an approach it calls deliberative alignment. confirmedas of 2024-12-21
- OpenAI o1 System Card · arXiv (OpenAI) · 2024-12-21 (retrieved 2026-10-10)
- [12]
Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03
- Reasoning models don't always say what they think · Anthropic (retrieved 2026-10-10)
- [13]
In December 2024 the ARC Prize Foundation reported that a preview of OpenAI's o3 scored 75.7% on the ARC-AGI-1 semi-private set within its $10,000 compute limit, and 87.5% in a configuration using about 172 times more compute. confirmedas of 2024-12-20
- OpenAI o3 Breakthrough High Score on ARC-AGI-Pub · ARC Prize Foundation · 2024-12-20 (retrieved 2026-10-10)
- [14]
The ARC Prize Foundation argued after o1's release that when AI systems may use a variable amount of test-time compute, no single benchmark score is objective, because accuracy depends on the compute allowed. confirmedas of 2024-09-13
- OpenAI o1 Results on ARC-AGI-Pub · ARC Prize Foundation · 2024-09-13 (retrieved 2026-10-10)
- [15]
Under Time Horizon 1.1, METR estimated 50% time horizons of about 320 minutes for Claude Opus 4.5, 214 minutes for GPT-5 and 121 minutes for o3. confirmedas of 2026-01-29
- Time Horizon 1.1 · METR · 2026-01-29 (retrieved 2026-10-10)
- [16]
OpenAI rated GPT-6 Astra at a Critical level of cybersecurity capability, meaning that with tools and access it can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [17]
DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Abstract (retrieved 2026-10-10)
- [18]
The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning · arXiv (DeepSeek-AI); published in Nature 645 · 2025-01-22 · Submission history and journal reference (retrieved 2026-10-10)
- [19]
Google DeepMind described Deep Think as exploring multiple solution paths in parallel and said the IMO model was trained with reinforcement learning techniques focused on multi-step problem solving. confirmedas of 2025-07-21
- Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad · Google DeepMind · Technical approach section (retrieved 2026-10-10)
- [20]
OpenAI's GPT-5.6, released with a system card on 9 July 2026, is a family of three models, Sol, Terra and Luna, which OpenAI treats as High capability in cybersecurity and in biological and chemical risk. confirmedas of 2026-07-09
- GPT-5.6 System Card · OpenAI · 2026-07-09 (retrieved 2026-10-10)
- [21]
OpenAI reported that GPT-5.6 Sol and Terra could find vulnerabilities and pieces of exploits but could not carry out autonomous end-to-end attacks against hardened targets in its testing. confirmedas of 2026-07-09
- GPT-5.6 System Card · OpenAI · 2026-07-09 (retrieved 2026-10-10)
- [22]
OpenAI's deployment safety hub lists system cards for GPT-5.6 (9 July 2026), GPT-6 Astra (3 September 2026), a GPT-6.1 Sol addendum (29 September 2026) and GPT-6 Sol and Luna (7 October 2026). confirmedas of 2026-10-10
- OpenAI Deployment Safety Hub · OpenAI · System card list (retrieved 2026-10-10)
- [23]
OpenAI's 7 October 2026 system card says GPT-6 Sol and GPT-6 Luna replace GPT-5.6 models in ChatGPT and are rated High, but below Critical, in cybersecurity and biological and chemical capability. confirmedas of 2026-10-07
- GPT-6 Sol and GPT-6 Luna: October 2026 update · OpenAI Deployment Safety Hub · 2026-10-07 (retrieved 2026-10-10)
- [24]
OpenAI said neither GPT-6 Sol nor GPT-6 Luna reaches its High threshold for AI self-improvement. confirmedas of 2026-10-07
- GPT-6 Sol and GPT-6 Luna: October 2026 update · OpenAI Deployment Safety Hub · 2026-10-07 (retrieved 2026-10-10)
- [25]
On 29 September 2026 OpenAI introduced GPT-6.1 Sol, which it says has capabilities comparable to GPT-6 Astra and treats as Critical in cybersecurity and High in biological and chemical capability. confirmedas of 2026-09-29
- Addendum to GPT-6 Astra System Card: GPT-6.1 Sol · OpenAI · 2026-09-29 (retrieved 2026-10-10)
- [26]
OpenAI said its safeguards for GPT-6 Astra include universal monitoring of full trajectories, including chains of thought, across tool-using inference in its external deployment. confirmedas of 2026-09-03
- GPT-6 Astra system card · OpenAI Deployment Safety Hub · 2026-09-03 (retrieved 2026-10-10)
- [27]
On 3 September 2026 the ARC Prize Foundation reported that OpenAI's GPT-6 Astra scored 62.7% on the ARC-AGI-3 semi-private set with its standard harness, at a cost of $26,098, and 99.9% with a provider-adapted harness, calling the result a noticeable step-function change in frontier model capabilities. confirmedas of 2026-09-03
- OpenAI's GPT-6 Astra on ARC-AGI-3 · ARC Prize Foundation · 2026-09-03 (retrieved 2026-10-10)
- [28]
The ARC Prize Foundation launched ARC-AGI-3 on 25 March 2026, reporting that humans scored 100% while frontier AI scored 0.51%. confirmedas of 2026-03-25
- Announcing ARC-AGI-3 · ARC Prize Foundation · 2026-03-25 · Announcement summary (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"OpenAI o1." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/openai-o1
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow reasoning models workHow AI reasoning models think step by step: chain of thought, reinforcement learning on checkable tasks and test-time compute.
- ExplainerFrontier AI in 2026: a crash courseA crash course on frontier AI in 2026: how reasoning models and AI agents work, who builds them, and where the frontier stands now.
- AnalysisDo reasoning models really reason? The debate over their limitsReasoning models win maths olympiads yet fail some simple tasks, and their written reasoning is not always faithful. The evidence, weighed.
- WikiARC-AGIARC-AGI is a benchmark series of tasks easy for people and hard for AI. ARC-AGI-3 went from 0.51% to 62.7% for AI within six months.
- WikiClaude MythosClaude Mythos is Anthropic's most capable model class, first released as a gated preview for cyber defence and later as Claude Fable.
- WikiDeepSeek-R1DeepSeek-R1 reported that reinforcement learning alone can teach a language model to reason. Its findings, publication and successors.