Skip to content
ContentLora

    Tip: press / anywhere to search.

    product

    DeepSeek-R1

    Also known as DeepSeek R1, R1, DeepSeek-R1-Zero

    DeepSeek-R1 is a reasoning model from the AI company DeepSeek whose paper reported that reasoning can be developed through pure reinforcement learning, without human-labelled reasoning examples.[1] First posted in January 2025, the work was later published in Nature.[2] DeepSeek has since folded reasoning into its V4 models, released with open weights in April 2026.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligenceScience
    Key facts

    DeepSeek-R1 is a publicly documented and peer-reviewed account of how to train a reasoning model.[1][2] Its central claim is that reinforcement learning alone, without human-labelled reasoning trajectories, can develop reasoning in a language model.[1]

    The research

    The paper, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”, was first posted on arXiv on 22 January 2025, about a month after OpenAI’s openai-o1 system card.[2][4] It reported that pure RL made human-written reasoning demonstrations unnecessary.[1] During training the model developed behaviours such as self-reflection, verification and dynamic strategy adaptation.[5] The authors reported that the RL-trained models beat conventionally supervised ones on verifiable tasks, such as mathematics, coding competitions and STEM problems, where answers can be checked automatically.[6] The work was later peer reviewed and published in Nature, volume 645.[2]

    Release

    DeepSeek released R1 on 20 January 2025, claiming performance on par with OpenAI o1, and put the code and models under the MIT License.[7] It also open-sourced six smaller models distilled from R1.[8]

    How it was trained

    The first model, R1-Zero, was trained with reinforcement learning using the GRPO algorithm on top of DeepSeek-V3 Base. Its reward depended only on whether the final answer was correct; the reasoning itself was not constrained.[9] As training went on, its average pass@1 score on the AIME 2024 maths competition rose from 15.6% to 77.9%.[10] The model also learned by itself to think for longer, producing longer responses as training progressed, a link between training and test-time-compute.[11] R1-Zero had flaws: its reasoning was hard to read and sometimes mixed English and Chinese, which the final R1 training pipeline was designed to fix.[12]

    Faithfulness of its reasoning

    Because R1 shows its chain of thought, researchers have used it to test whether written reasoning reflects what a model actually does. In an Anthropic study, DeepSeek R1 mentioned a planted hint that it used in 39% of cases, against 25% for Claude 3.7 Sonnet.[13] That gap between stated and actual reasoning is part of the wider debate on reasoning models’ limits.[14]

    From R1 to V4

    DeepSeek’s newer models put reasoning and general use in one model. On 24 April 2026 it announced V4 preview models offering both a thinking and a non-thinking mode, with weights released on Hugging Face.[3] V4-Pro has 1.6 trillion total parameters with 49 billion active per token, and V4-Flash has 284 billion with 13 billion active; both support a 1-million-token context.[15] DeepSeek said its older deepseek-chat and deepseek-reasoner API models would be retired after 24 July 2026.[16] V4-Pro became generally available on 13 August 2026 with selectable low, high and max reasoning effort.[17] On 10 September DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model, and said it was phasing out V4-Pro because tests put the new model ahead.[18]

    Place in the field

    The 2026 AI Index reports that US and Chinese models have traded places at the top of performance rankings several times since early 2025, and that in March 2026 Anthropic’s top model led by just 2.7%.[19] Because DeepSeek publishes V4’s weights on Hugging Face, outside researchers can download and run the model themselves.[3]

    The US government’s CAISI evaluated V4 Pro in April 2026 and estimated that it lagged the US frontier by about eight months: on CAISI’s tests it performed like GPT-5, although DeepSeek’s self-reported results compared it with newer models.[20] CAISI also found V4 more cost-efficient than models of similar capability.[21] An earlier CAISI study, in September 2025, found DeepSeek’s R1-0528 12 times more likely than the US frontier models it tested to follow malicious instructions.[22]

    Questions readers ask

    What did DeepSeek-R1 show?

    Its paper reported that a language model's reasoning can be developed through pure reinforcement learning, without people writing example reasoning for it to copy.[1]

    Was the DeepSeek-R1 research peer reviewed?

    Yes. After first appearing on arXiv in January 2025, it was published in Nature, volume 645, in 2025.[2]

    What replaced DeepSeek-R1?

    DeepSeek's V4 models, announced in April 2026, combine thinking and non-thinking modes, and DeepSeek retired its older deepseek-reasoner API model after 24 July 2026.[3][16]

    How large is DeepSeek V4?

    V4-Pro has 1.6 trillion total and 49 billion active parameters, and V4-Flash has 284 billion total and 13 billion active, both with a 1-million-token context.[15]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04

    2. [2]

      The DeepSeek-R1 paper was first posted on arXiv on 22 January 2025 and was later published in Nature (volume 645, 2025). confirmedas of 2026-01-04

    3. [3]

      DeepSeek V4 offers both a thinking and a non-thinking mode, and its weights were released openly on Hugging Face. confirmedas of 2026-04-24

    4. [4]

      OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21

    5. [5]

      According to the DeepSeek-R1 paper, reinforcement learning led the model to develop behaviours such as self-reflection, verification and dynamic strategy adaptation. confirmedas of 2026-01-04

    6. [6]

      The DeepSeek-R1 paper reports that its reinforcement-learning-trained models did better than conventionally supervised models on verifiable tasks such as mathematics, coding competitions and STEM problems. confirmedas of 2026-01-04

    7. [7]

      DeepSeek released DeepSeek-R1 on 20 January 2025 under the MIT License, saying it performed on par with OpenAI o1. confirmedas of 2025-01-20

    8. [8]

      Alongside R1, DeepSeek open-sourced six smaller models distilled from it. confirmedas of 2025-01-20

    9. [9]

      In DeepSeek-R1-Zero's reinforcement learning, which used the GRPO algorithm on top of DeepSeek-V3 Base, the reward depended only on whether final answers were correct, not on the reasoning process. confirmedas of 2025-09-17

    10. [10]

      During reinforcement learning, DeepSeek-R1-Zero's average pass@1 score on the AIME 2024 maths benchmark rose from 15.6% to 77.9%. confirmedas of 2025-09-17

    11. [11]

      The DeepSeek-R1 paper reports that R1-Zero learned to spend more thinking time, producing longer responses, as reinforcement learning progressed. confirmedas of 2025-09-17

    12. [12]

      DeepSeek reported that R1-Zero's reasoning had poor readability and sometimes mixed English and Chinese, problems the final R1 pipeline was built to address. confirmedas of 2025-09-17

    13. [13]

      Anthropic researchers found that when given hints, Claude 3.7 Sonnet mentioned the hint in its reasoning 25% of the time and DeepSeek R1 39% of the time. confirmedas of 2025-04-03

    14. [14]

      A 2025 paper by more than 40 researchers argued that monitoring reasoning models' chains of thought for intent to misbehave is a valuable safety opportunity, but that chain-of-thought monitorability may be fragile. confirmedas of 2025-12-07

    15. [15]

      DeepSeek announced V4 preview models on 24 April 2026, V4-Pro with 1.6 trillion total and 49 billion active parameters and V4-Flash with 284 billion total and 13 billion active, both with a 1-million-token context. confirmedas of 2026-04-24

    16. [16]

      DeepSeek said its older deepseek-chat and deepseek-reasoner API models would be retired after 24 July 2026. confirmedas of 2026-04-24

    17. [17]

      On 13 August 2026 DeepSeek made V4-Pro generally available, with selectable low, high and max reasoning effort. confirmedas of 2026-08-13

    18. [18]

      On 10 September 2026 DeepSeek released V4.1-Flash, a 552-billion-parameter mixture-of-experts model, and said it was phasing out V4-Pro because tests put V4.1-Flash ahead of it. confirmedas of 2026-09-10

    19. [19]

      The 2026 AI Index reports that US and Chinese models have traded places at the top of performance rankings several times since early 2025, and that as of March 2026 Anthropic's top model led by just 2.7%. confirmedas of 2026-03-31

    20. [20]

      The US Center for AI Standards and Innovation (CAISI) evaluated DeepSeek V4 Pro in April 2026 and estimated that its capabilities lagged the US frontier by about eight months, performing on CAISI's tests similarly to GPT-5 even though DeepSeek's self-reported results compared it with newer models. confirmedas of 2026-05-01

    21. [21]

      CAISI found DeepSeek V4 more cost-efficient than other models of similar capability. confirmedas of 2026-05-01

    22. [22]

      In September 2025 CAISI published an evaluation of three DeepSeek models against four US models across 19 benchmarks, finding the DeepSeek R1-0528 model 12 times more likely to follow malicious instructions than the US frontier models evaluated. confirmedas of 2025-09-30

    Revision history (2)
    1. Page created.
    2. Added the January 2025 release and licence, how R1-Zero was trained and what it learned, R1-Zero's readability problems, DeepSeek V4-Pro GA and V4.1-Flash, and CAISI's evaluation of DeepSeek V4.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "DeepSeek-R1." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/deepseek-r1

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.