Skip to content
ContentLora

    Tip: press / anywhere to search.

    Explainer

    How AI agents work

    An AI agent is a system in which a language model directs its own steps and tool use to finish a task,[1] typically by calling tools, checking the results and repeating in a loop.[2] Agents went from about 12% to about 66% success on a benchmark of real computer tasks, according to the 2026 AI Index.[3]

    Editor reviewedUpdated Frontier AIArtificial intelligenceComputing

    An AI agent is a language model that does not just answer, but acts: it plans, uses tools such as code execution or a web browser, looks at what happened and decides what to do next.[2] Agents are where frontier AI meets real work, and their benchmark scores have risen sharply.[3][4]

    The agent loop

    Think of a chatbot that can also press buttons. You give it a goal. It decides on a step, such as running a search or editing a file, sees the result, and decides the next step, until it is done. Anthropic describes agents as language models using tools based on feedback from their environment in a loop.[2]

    Anthropic separates workflows, where language models and tools follow predefined code paths, from agents, where the model dynamically directs its own process and tool use.[5][1] The building block is an “augmented LLM” with retrieval, tools and memory.[6] An early example of the pattern is ReAct (2022), which interleaved reasoning traces with actions such as API calls so that observations could guide further reasoning.[7] METR attributes longer agent task horizons mainly to better reliability, mistake recovery, logical reasoning and tool use.[8]

    Using a computer like a person

    Some agents operate an ordinary computer through screenshots, moving the mouse and typing. Anthropic released this “computer use” ability in October 2024 and called it experimental and error-prone.[9] On OSWorld, a test of real computer tasks, humans succeeded about 72% of the time in 2024 while the best AI managed about 12%.[10] By the 2026 AI Index, agents had reached about 66%.[3]

    OSWorld comprises 369 tasks across operating systems; at release, humans scored 72.36% and the best model 12.24%.[10] Claude 3.5 Sonnet’s computer-use beta scored 14.9% screenshot-only in October 2024.[11] The 2026 AI Index reports agent success of about 66% on OSWorld,[3] and Anthropic reported 81.8% partial accuracy for Claude Opus 5.5 on OSWorld 2.1 in September 2026.[12]

    Connecting to tools and data

    An agent is only as useful as the tools it can reach. The model-context-protocol (MCP) is an open standard for plugging AI apps into files, databases and services, which its documentation compares to a USB-C port.[13] It is now supported by assistants such as Claude and ChatGPT and by coding tools such as Visual Studio Code and Cursor.[14]

    MCP, introduced by Anthropic in November 2024, standardises how AI applications connect to data sources, tools and prompt workflows.[15][13] In December 2025 it moved to the Linux Foundation’s new Agentic AI Foundation alongside OpenAI’s AGENTS.md and Block’s goose.[16] Connecting agents to external content also creates attack surface: OpenAI reports prompt-injection robustness figures, such as 99.79% for GPT-6 Astra on indirect attacks, in its system cards.[17]

    How long agents can work

    A useful way to track agents is to ask how long a task, measured in human working time, they can finish on their own. In early 2025 the best models managed tasks of about 50 minutes half the time.[18] By March 2026, the research group METR estimated at least 16 hours for an early version of Claude Mythos Preview.[19]

    METR’s 50% time horizon measures the human-expert duration of tasks a model completes with 50% success.[20] It doubled about every seven months from 2019 to early 2025,[18] and METR’s January 2026 revision estimated about 89 days since 2024.[21] Measurements above 16 hours are unreliable with its current task suite.[19] Coding agents are also tracked on swe-bench, where the 2026 AI Index reports a jump from 60% to near 100% on the Verified subset in one year.[22]

    Benchmarks versus real work

    Benchmark gains do not translate automatically into productivity. In METR’s 2025 trial, experienced developers using early-2025 AI tools took 19% longer on tasks.[23] Its 2026 follow-up said true gains were likely much higher but that its evidence was very weak because of selection effects.[24] The analysis page on the pace of agent progress weighs these signals.[25]

    Questions readers ask

    What is the difference between an AI agent and an AI workflow?

    In Anthropic's definitions, a workflow orchestrates a language model and tools through predefined code paths, while an agent lets the model dynamically direct its own process and tool use.[5][1]

    Can AI agents use a normal computer?

    Yes. Anthropic released computer use in public beta in October 2024, letting Claude look at a screen, move a cursor, click and type. On the OSWorld benchmark, agents have gone from about 12% to about 66% task success.[9][3]

    How do agents connect to apps and data?

    Increasingly through the Model Context Protocol, an open standard that the MCP documentation compares to a USB-C port for AI applications.[13]

    How long can AI agents work on their own?

    METR estimated that an early version of Claude Mythos Preview had a 50% time horizon of at least 16 hours on software tasks in March 2026, at the edge of what METR can measure.[19]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      Anthropic defines agents as systems in which language models dynamically direct their own processes and tool use, keeping control over how they accomplish a task. confirmedas of 2024-12-19

    2. [2]

      Anthropic describes agents as typically language models using tools based on feedback from their environment in a loop, checking results such as tool outputs or code execution at each step. confirmedas of 2024-12-19

    3. [3]

      The 2026 AI Index reports that AI agents went from about 12% to about 66% task success on OSWorld. confirmedas of 2026-04-01

    4. [4]

      The 2026 AI Index reports that agent success on a benchmark of real-world tasks rose from 20% to 77.3%. confirmedas of 2026-04-01

    5. [5]

      Anthropic distinguishes agents from workflows, which it defines as systems where language models and tools are orchestrated through predefined code paths. confirmedas of 2024-12-19

    6. [6]

      Anthropic describes the basic building block of agentic systems as a language model augmented with retrieval, tools and memory. confirmedas of 2024-12-19

    7. [7]

      The 2022 ReAct method has language models generate reasoning traces and task-specific actions in an interleaved way, so actions such as querying a Wikipedia API can gather information that guides further reasoning. confirmedas of 2022-10-06

    8. [8]

      METR attributed the growth in time horizons mainly to better reliability, ability to recover from mistakes, logical reasoning and tool use. confirmedas of 2025-03-18

    9. [9]

      On 22 October 2024 Anthropic released computer use in public beta, letting Claude operate a computer by looking at a screen, moving a cursor, clicking and typing, and described the feature as experimental and error-prone. confirmedas of 2024-10-22

    10. [10]

      OSWorld, a 2024 benchmark of 369 real computer tasks across operating systems, found humans succeeded on 72.36% of tasks against 12.24% for the best AI model at the time. confirmedas of 2024-04-11

    11. [11]

      At launch Anthropic reported that Claude 3.5 Sonnet scored 14.9% on OSWorld in screenshot-only mode, against 7.8% for the next-best system. confirmedas of 2024-10-22

    12. [12]

      Anthropic reported that Claude Opus 5.5 scored 66.4% on Terminal-Bench 4.0, 81.8% on OSWorld 2.1 and 67.7% on Humanity's Last Exam with tools. confirmedas of 2026-09-22

    13. [13]

      The MCP documentation describes MCP as an open-source standard for connecting AI applications to data sources, tools and workflows, comparing it to a USB-C port for AI applications. confirmedas of 2026-10-10

    14. [14]

      As of October 2026 MCP is supported by AI assistants including Claude and ChatGPT and by developer tools including Visual Studio Code and Cursor. confirmedas of 2026-10-10

    15. [15]

      Anthropic introduced the Model Context Protocol (MCP) on 25 November 2024 as an open standard for connecting AI assistants to the systems where data lives, such as content repositories, business tools and development environments. confirmedas of 2024-11-25

    16. [16]

      On 9 December 2025 the Linux Foundation formed the Agentic AI Foundation (AAIF), with Anthropic's MCP, Block's goose and OpenAI's AGENTS.md as founding projects. confirmedas of 2025-12-09

    17. [17]

      OpenAI reported 99.79% robustness for GPT-6 Astra against indirect prompt-injection attacks in its evaluations. confirmedas of 2026-09-03

    18. [18]

      METR's 2025 study found that the 50% time horizon of frontier AI models had doubled approximately every seven months since 2019, with Claude 3.7 Sonnet at around 50 minutes. confirmedas of 2025-03-18

    19. [19]

      METR estimated that an early version of Claude Mythos Preview, evaluated in March 2026, had a 50% time horizon of at least 16 hours (95% confidence interval 8.5 to 55 hours), at the upper end of what its task suite can measure. confirmedas of 2026-05-10

    20. [20]

      METR's 50%-task-completion time horizon is the length of tasks, measured by how long human experts take, that an AI model completes with 50% success. confirmedas of 2025-03-18

    21. [21]

      In January 2026 METR released Time Horizon 1.1, expanding its task suite from 170 to 228 tasks and its tasks of eight hours or more from 14 to 31, and estimating a doubling time of about 131 days since 2023 and about 89 days since 2024. confirmedas of 2026-01-29

    22. [22]

      The 2026 AI Index reports that performance on the SWE-bench Verified coding benchmark rose from 60% to near 100% in a single year. confirmedas of 2026-04-01

    23. [23]

      In a 2025 randomized controlled trial with 16 experienced open-source developers on 246 tasks, METR found that allowing early-2025 AI tools increased task completion time by 19%, although developers had expected a 24% speedup. confirmedas of 2025-07-12

    24. [24]

      METR said the true productivity gains from AI tools were likely much higher than its late-2025 study measured, but that its data provided only very weak evidence, and that it would change its experiment design. confirmedas of 2026-02-24

    25. [25]

      METR's follow-up study with 57 developers and more than 800 tasks estimated a speedup of -18% (CI -38% to +9%) for returning developers and -4% (CI -15% to +9%) for new ones, results it said were biased by selection effects. confirmedas of 2026-02-24

    Revision history (1)
    1. Page created.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "How AI agents work." ContentLora, updated Oct 10, 2026. https://contentlora.com/explain/how-ai-agents-work

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.