Skip to content
ContentLora

    Tip: press / anywhere to search.

    organization

    Protein Data Bank (PDB)

    Also known as PDB, wwPDB, Worldwide Protein Data Bank, RCSB PDB

    The Protein Data Bank is the single global archive of experimentally determined 3D structures of proteins, nucleic acids and complexes, run since 1971 and open to all without usage limits.[1][2] It is the training data behind protein AI: AlphaFold2 was trained on a copy of it.[3]

    Editor reviewedUpdated AI for scienceLife sciencesScience
    Key facts

    Protein AI is built on the Protein Data Bank (PDB). The PDB has collected experimentally determined 3D structures of proteins, nucleic acids and complex assemblies since 1971, as the single archive for the field.[1] Its data are public domain with no limits on use, which made it possible to train models such as alphafold on decades of shared work.[2][3]

    What it holds

    As of 10 October 2026 the RCSB PDB website listed 260,626 experimental structures.[4] Before AlphaFold2, these experiments had covered about 100,000 unique proteins, a small share of the billions of known sequences, because each structure could take months or years.[5]

    Who runs it

    The archive is managed by the Worldwide PDB (wwPDB) partnership, which promises universal open access to the data.[1][2] Regional sites such as the RCSB PDB provide search and statistics on top of the shared archive.[4]

    Role in AI for science

    AlphaFold2’s developers trained their model on a copy of the PDB downloaded on 28 August 2019.[3] That dependence explains why experimental structural biology remains central even after AI: new experimental structures are both training data and the test set for blind assessments such as CASP.[6]

    Experimental and predicted side by side

    The RCSB PDB website now also indexes 1,062,058 computed structure models from AlphaFold DB and ModelArchive, displayed separately from experimental entries.[4] Keeping the two apart matters: a predicted model is a hypothesis, while an experimental structure is a measurement with its own error estimates. AlphaFold DB alone added 1.7 million predicted homodimers in March 2026.[7] By October 2026 the AlphaFold Database said it offered over 260 million predictions, roughly a thousand times the number of experimental entries in the PDB.[8][4]

    Blind tests depend on it

    The CASP experiment, which first showed AlphaFold2’s accuracy, works by asking predictors to model structures that are about to be solved experimentally, before they appear in public archives.[9] Its 17th round closed predictions on 11 September 2026, with results due on 30 November 2026.[6] CASP17 concentrates on cases where deep learning has yet to deliver, such as immune complexes, protein–ligand complexes, nucleic acids and proteins that switch between shapes.[10] Each of those categories needs fresh experimental structures, which is why new experimental deposits remain essential.[4]

    The next generation of training data

    Open efforts to retrain AlphaFold-class models also start from public structural data. The OpenFold project, for example, publishes the full training data needed to reproduce its OpenFold3-preview model, an open reimplementation of AlphaFold 3.[11][12] The archive also defines what counts as a hard design test. When Chai Discovery evaluated its Chai-2 antibody model in 2025, it chose 52 targets that had no existing antibody or nanobody binder in the PDB.[13]

    Public funding is also going into new structural data. One of the US Genesis Mission’s October 2026 Phase II awards went to a UC San Diego project that plans to expand the RNA structure database five-fold to train AI models.[14] See genesis-mission.

    Questions readers ask

    How many structures are in the PDB?

    As of 10 October 2026 the RCSB PDB listed 260,626 experimental structures.[4]

    Does the PDB include AI predictions?

    AI predictions are kept apart from experimental entries; the RCSB PDB site indexes over a million computed structure models from AlphaFold DB and ModelArchive.[4]

    Why does the PDB matter for AI?

    Protein structure models need many solved examples to learn from; AlphaFold2 was trained on a PDB snapshot from August 2019.[3]

    Sources

    Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.

    1. [1]

      Since 1971 the Protein Data Bank has been the single repository of 3D structures of proteins, nucleic acids and complex assemblies, managed by the Worldwide PDB partnership. confirmedas of 2026-10-10

    2. [2]

      The wwPDB provides open access to its structural biology data with no limitations on usage. confirmedas of 2026-10-10

    3. [3]

      AlphaFold2 was trained on a copy of the Protein Data Bank downloaded in August 2019. confirmedas of 2021-07-15

    4. [4]

      As of 10 October 2026 the RCSB PDB listed 260,626 experimental structures and 1,062,058 computed structure models from AlphaFold DB and ModelArchive. confirmedas of 2026-10-10

    5. [5]

      Before AlphaFold2, experiments had determined structures for around 100,000 unique proteins, a small fraction of the billions of known protein sequences, because each structure takes months to years of work. confirmedas of 2021-07-15

    6. [6]

      The CASP17 prediction season ended on 11 September 2026, with assessment results due on 30 November 2026, a day before the CASP17 conference in Rome. confirmedas of 2026-10-10

      • CASP17 home page · Protein Structure Prediction Center (UC Davis) (retrieved 2026-10-10)
    7. [7]

      In March 2026 EMBL-EBI, Google DeepMind, NVIDIA and Seoul National University added 1.7 million high-confidence predicted homodimer structures to the AlphaFold Database. confirmedas of 2026-03-16

    8. [8]

      As of October 2026 the AlphaFold Protein Structure Database, developed by Google DeepMind and EMBL-EBI, said it provides open access to over 260 million protein structure predictions. confirmedas of 2026-10-10

    9. [9]

      AlphaFold2, published in Nature in July 2021, was validated in the CASP14 blind assessment and predicted structures with accuracy competitive with experiment in a majority of cases. confirmedas of 2021-07-15

    10. [10]

      CASP17 emphasises areas where deep learning has yet to deliver, including immune complexes, protein–ligand complexes, nucleic acids and conformational ensembles. confirmedas of 2026-10-10

      • CASP17 home page · Protein Structure Prediction Center (UC Davis) (retrieved 2026-10-10)
    11. [11]

      The OpenFold project provides the full training data needed to reproduce OpenFold3-preview, including a 13-million-sequence distillation dataset modelled on the one described in the AlphaFold 3 paper. confirmedas of 2026-10-10

    12. [12]

      OpenFold3-preview, from Columbia University's AlQuraishi Lab and the OpenFold consortium, aims to be a bitwise reproduction of AlphaFold 3 and is available under the Apache 2.0 licence for academic and commercial use. confirmedas of 2026-10-10

    13. [13]

      In a preprint posted in July 2025, Chai Discovery reported that its Chai-2 model achieved a 16% hit rate in fully de novo antibody design and found at least one binder for 50% of 52 targets, testing 20 or fewer designs per target, none of which had an existing antibody binder in the Protein Data Bank. confirmedas of 2025-07-06

    14. [14]

      DOE's October 2026 Phase II Genesis Mission awards include a Commonwealth Fusion Systems digital twin of a fusion demonstration device, University of Washington enzyme design tools, a UC San Diego project to expand the RNA structure database five-fold and a Fermilab project to design rugged microchips for extreme environments. confirmedas of 2026-10-08

    Revision history (2)
    1. Page created.
    2. Added the scale gap between the PDB and the AlphaFold Database, open retraining efforts, the PDB's role in testing antibody design, and the Genesis Mission RNA-structure project.

    Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.

    Cite this page

    "Protein Data Bank (PDB)." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/protein-data-bank

    Spotted an error? Suggest a correction or emailcorrections@contentlora.com.