Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    How large language models work

    A large language model is a neural network trained on huge amounts of text to predict what comes next. Most are built on the Transformer, an architecture based on attention,[1] and they improve predictably as models, data and compute grow.[2] Extra training with human feedback turns a raw text predictor into a helpful assistant.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligenceComputing

    Large language models (LLMs) are the base layer of everything else in frontier AI: reasoning models and AI agents are LLMs with extra training and extra tools. This page explains the three ideas you need first, which are the Transformer, scaling and training with human feedback.[1][2][3]

    Predicting the next word

    At heart, a language model is a very large pattern-matcher for text. GPT-3, a landmark 2020 model with 175 billion adjustable numbers called parameters, could do new tasks just from a written prompt with a few examples, without being retrained.[4] Its makers found that making models bigger greatly improved this ability.[4]

    An LLM is trained to predict the next token in a sequence. The GPT-3 paper showed that scaling such a model to 175 billion parameters greatly improved task-agnostic few-shot performance: tasks are specified in the prompt, with no gradient updates.[4] This “in-context learning” is why one model can translate, answer questions and write code without separate task-specific training.[4]

    The Transformer

    Almost all modern LLMs use a design called the Transformer, published in 2017.[1] Its key trick is “attention”, which lets the model weigh how much every word in the text matters to every other word. The Transformer relies on attention alone and drops older designs that read text strictly in order.[1]

    The Transformer is based solely on attention mechanisms and dispenses with recurrence and convolutions entirely.[1] In practice, each position in the input attends to every other position through learned weights, and stacks of these attention layers form the model.[1]

    Scaling laws

    In 2020, researchers found a simple pattern: as you make a model bigger, give it more data and spend more computing power on training, its error falls smoothly and predictably.[2] That predictability is why AI companies have spent ever more on training runs.[5]

    Kaplan et al. found that loss scales as a power law with model size, dataset size and training compute, with some trends spanning more than seven orders of magnitude.[2] Since 2024, labs have added a second axis: spending more compute at inference time, which reasoning models exploit (see test-time-compute).[6] Global corporate AI investment reached $581.7 billion in 2025.[5]

    From text predictor to assistant

    A raw model trained only to predict text is not a good assistant. In 2022, the InstructGPT work showed a fix: people wrote example answers and ranked the model’s outputs, and the model was trained to produce the kind of answers people preferred.[3] A small model trained this way was preferred to a model more than 100 times its size.[7]

    InstructGPT used supervised fine-tuning on labeller demonstrations followed by reinforcement learning from human feedback (RLHF), with a reward signal learned from human rankings of model outputs.[3] The 1.3B InstructGPT model was preferred to the 175B GPT-3 in human evaluations.[7] Reasoning models extend this post-training stage with reinforcement learning on tasks whose answers can be checked automatically.[8][9]

    Where this leads

    LLMs remain uneven. The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time,[10] and the International AI Safety Report 2026 notes that systems can excel at some difficult tasks while failing at simpler ones.[11] The next pages cover how labs taught LLMs to reason step by step (openai-o1, deepseek-r1) and to act as agents (model-context-protocol).[8][12]

    Questions readers ask

    What architecture do most large language models use?

    Most use the Transformer, introduced in 2017, which is based on attention mechanisms and drops the recurrence and convolutions used in earlier networks.[1]

    Why do AI labs keep building bigger models?

    A 2020 study found that language-model loss falls as a power law as model size, data and compute grow, with some trends holding across more than seven orders of magnitude.[2]

    What is RLHF?

    Reinforcement learning from human feedback. In InstructGPT, people ranked model outputs and the model was then trained with reinforcement learning to produce outputs people prefer.[3]

    Does a bigger model always give better answers?

    Not always. In human evaluations, outputs from a 1.3-billion-parameter InstructGPT model were preferred to those of the 175-billion-parameter GPT-3.[7]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      The Transformer architecture, introduced in a 2017 paper, is based solely on attention mechanisms and dispenses with recurrence and convolutions. confirmedas of 2017-06-12

    2. [2]

      A 2020 study found that language-model loss falls as a power law as model size, dataset size and training compute grow, with some trends spanning more than seven orders of magnitude. confirmedas of 2020-01-23

    3. [3]

      InstructGPT was trained in two stages, supervised fine-tuning on human-written demonstrations and then reinforcement learning from human feedback (RLHF) using human rankings of model outputs. confirmedas of 2022-03-04

    4. [4]

      The 2020 GPT-3 paper reported that scaling up language models greatly improves few-shot performance, with the 175-billion-parameter model doing new tasks from text prompts alone, without gradient updates. confirmedas of 2020-05-28

    5. [5]

      The 2026 AI Index reports that global corporate AI investment reached $581.7 billion in 2025, up 130% from the year before. confirmedas of 2026-04-01

    6. [6]

      The International AI Safety Report 2026 describes inference-time scaling, in which models use more computing power to generate intermediate steps before giving a final answer, as a major way developers now improve capabilities. confirmedas of 2026-02-24

    7. [7]

      In human evaluations, outputs of the 1.3-billion-parameter InstructGPT model were preferred to those of the 175-billion-parameter GPT-3, despite it having 100 times fewer parameters. confirmedas of 2022-03-04

    8. [8]

      OpenAI's o1 model series is trained with large-scale reinforcement learning to reason using a chain of thought before answering. confirmedas of 2024-12-21

    9. [9]

      DeepSeek reported that reasoning abilities in its DeepSeek-R1 work could be developed through pure reinforcement learning, without human-labelled reasoning trajectories. confirmedas of 2026-01-04

    10. [10]

      The 2026 AI Index reports that the top model read analog clocks correctly only 50.1% of the time, an example of uneven capabilities. confirmedas of 2026-04-01

    11. [11]

      The International AI Safety Report 2026 says advanced AI systems may excel at some difficult tasks while failing at simpler ones, such as counting objects in an image. confirmedas of 2026-02-24

    12. [12]

      Anthropic defines agents as systems in which language models dynamically direct their own processes and tool use, keeping control over how they accomplish a task. confirmedas of 2024-12-19

    Revision history (1)
    1. Page created.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "How large language models work." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-large-language-models-work

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.