organization
Protein Data Bank (PDB)
Also known as PDB, wwPDB, Worldwide Protein Data Bank, RCSB PDB
The Protein Data Bank is the single global archive of experimentally determined 3D structures of proteins, nucleic acids and complexes, run since 1971 and open to all without usage limits.[1][2] It is the training data behind protein AI: AlphaFold2 was trained on a copy of it.[3]
Key facts
Protein AI is built on the Protein Data Bank (PDB). The PDB has collected experimentally determined 3D structures of proteins, nucleic acids and complex assemblies since 1971, as the single archive for the field.[1] Its data are public domain with no limits on use, which made it possible to train models such as alphafold on decades of shared work.[2][3]
What it holds
As of 10 October 2026 the RCSB PDB website listed 260,626 experimental structures.[4] Before AlphaFold2, these experiments had covered about 100,000 unique proteins, a small share of the billions of known sequences, because each structure could take months or years.[5]
Who runs it
The archive is managed by the Worldwide PDB (wwPDB) partnership, which promises universal open access to the data.[1][2] Regional sites such as the RCSB PDB provide search and statistics on top of the shared archive.[4]
Role in AI for science
AlphaFold2’s developers trained their model on a copy of the PDB downloaded on 28 August 2019.[3] That dependence explains why experimental structural biology remains central even after AI: new experimental structures are both training data and the test set for blind assessments such as CASP.[6]
Experimental and predicted side by side
The RCSB PDB website now also indexes 1,062,058 computed structure models from AlphaFold DB and ModelArchive, displayed separately from experimental entries.[4] Keeping the two apart matters: a predicted model is a hypothesis, while an experimental structure is a measurement with its own error estimates. AlphaFold DB alone added 1.7 million predicted homodimers in March 2026.[7] By October 2026 the AlphaFold Database said it offered over 260 million predictions, roughly a thousand times the number of experimental entries in the PDB.[8][4]
Blind tests depend on it
The CASP experiment, which first showed AlphaFold2’s accuracy, works by asking predictors to model structures that are about to be solved experimentally, before they appear in public archives.[9] Its 17th round closed predictions on 11 September 2026, with results due on 30 November 2026.[6] CASP17 concentrates on cases where deep learning has yet to deliver, such as immune complexes, protein–ligand complexes, nucleic acids and proteins that switch between shapes.[10] Each of those categories needs fresh experimental structures, which is why new experimental deposits remain essential.[4]
The next generation of training data
Open efforts to retrain AlphaFold-class models also start from public structural data. The OpenFold project, for example, publishes the full training data needed to reproduce its OpenFold3-preview model, an open reimplementation of AlphaFold 3.[11][12] The archive also defines what counts as a hard design test. When Chai Discovery evaluated its Chai-2 antibody model in 2025, it chose 52 targets that had no existing antibody or nanobody binder in the PDB.[13]
Public funding is also going into new structural data. One of the US Genesis Mission’s October 2026 Phase II awards went to a UC San Diego project that plans to expand the RNA structure database five-fold to train AI models.[14] See genesis-mission.
Questions readers ask
How many structures are in the PDB?
As of 10 October 2026 the RCSB PDB listed 260,626 experimental structures.[4]
Does the PDB include AI predictions?
AI predictions are kept apart from experimental entries; the RCSB PDB site indexes over a million computed structure models from AlphaFold DB and ModelArchive.[4]
Why does the PDB matter for AI?
Protein structure models need many solved examples to learn from; AlphaFold2 was trained on a PDB snapshot from August 2019.[3]
Sources
Each numbered claim is a statement we checked against the sources listed with it. Status shows how well established it is.
- [1]
Since 1971 the Protein Data Bank has been the single repository of 3D structures of proteins, nucleic acids and complex assemblies, managed by the Worldwide PDB partnership. confirmedas of 2026-10-10
- Worldwide Protein Data Bank (wwPDB) home page · wwPDB (retrieved 2026-10-10)
- [2]
The wwPDB provides open access to its structural biology data with no limitations on usage. confirmedas of 2026-10-10
- Worldwide Protein Data Bank (wwPDB) home page · wwPDB (retrieved 2026-10-10)
- [3]
AlphaFold2 was trained on a copy of the Protein Data Bank downloaded in August 2019. confirmedas of 2021-07-15
- Highly accurate protein structure prediction with AlphaFold · Nature · 2021-07-15 · Methods (retrieved 2026-10-10)
- [4]
As of 10 October 2026 the RCSB PDB listed 260,626 experimental structures and 1,062,058 computed structure models from AlphaFold DB and ModelArchive. confirmedas of 2026-10-10
- RCSB Protein Data Bank home page (data statistics panel) · RCSB PDB (retrieved 2026-10-10)
- [5]
Before AlphaFold2, experiments had determined structures for around 100,000 unique proteins, a small fraction of the billions of known protein sequences, because each structure takes months to years of work. confirmedas of 2021-07-15
- Highly accurate protein structure prediction with AlphaFold · Nature · 2021-07-15 (retrieved 2026-10-10)
- [6]
The CASP17 prediction season ended on 11 September 2026, with assessment results due on 30 November 2026, a day before the CASP17 conference in Rome. confirmedas of 2026-10-10
- CASP17 home page · Protein Structure Prediction Center (UC Davis) (retrieved 2026-10-10)
- [7]
In March 2026 EMBL-EBI, Google DeepMind, NVIDIA and Seoul National University added 1.7 million high-confidence predicted homodimer structures to the AlphaFold Database. confirmedas of 2026-03-16
- Millions of protein complexes added to AlphaFold Database shed light on how proteins interact · EMBL · 2026-03-16 (retrieved 2026-10-10)
- [8]
As of October 2026 the AlphaFold Protein Structure Database, developed by Google DeepMind and EMBL-EBI, said it provides open access to over 260 million protein structure predictions. confirmedas of 2026-10-10
- AlphaFold Protein Structure Database · EMBL-EBI (retrieved 2026-10-10)
- [9]
AlphaFold2, published in Nature in July 2021, was validated in the CASP14 blind assessment and predicted structures with accuracy competitive with experiment in a majority of cases. confirmedas of 2021-07-15
- Highly accurate protein structure prediction with AlphaFold · Nature · 2021-07-15 (retrieved 2026-10-10)
- [10]
CASP17 emphasises areas where deep learning has yet to deliver, including immune complexes, protein–ligand complexes, nucleic acids and conformational ensembles. confirmedas of 2026-10-10
- CASP17 home page · Protein Structure Prediction Center (UC Davis) (retrieved 2026-10-10)
- [11]
The OpenFold project provides the full training data needed to reproduce OpenFold3-preview, including a 13-million-sequence distillation dataset modelled on the one described in the AlphaFold 3 paper. confirmedas of 2026-10-10
- aqlaboratory/openfold-3: A fully open source biomolecular structure prediction model based on AlphaFold3 (README) · AlQuraishi Lab and OpenFold consortium (GitHub) (retrieved 2026-10-10)
- [12]
OpenFold3-preview, from Columbia University's AlQuraishi Lab and the OpenFold consortium, aims to be a bitwise reproduction of AlphaFold 3 and is available under the Apache 2.0 licence for academic and commercial use. confirmedas of 2026-10-10
- aqlaboratory/openfold-3: A fully open source biomolecular structure prediction model based on AlphaFold3 (README) · AlQuraishi Lab and OpenFold consortium (GitHub) (retrieved 2026-10-10)
- [13]
In a preprint posted in July 2025, Chai Discovery reported that its Chai-2 model achieved a 16% hit rate in fully de novo antibody design and found at least one binder for 50% of 52 targets, testing 20 or fewer designs per target, none of which had an existing antibody binder in the Protein Data Bank. confirmedas of 2025-07-06
- Zero-shot antibody design in a 24-well plate · bioRxiv (Chai Discovery) · 2025-07-06 (retrieved 2026-10-10)
- [14]
DOE's October 2026 Phase II Genesis Mission awards include a Commonwealth Fusion Systems digital twin of a fusion demonstration device, University of Washington enzyme design tools, a UC San Diego project to expand the RNA structure database five-fold and a Fermilab project to design rugged microchips for extreme environments. confirmedas of 2026-10-08
- Energy Department Announces New Genesis Mission Awards to Advance Super Intelligence for Science · U.S. Department of Energy · 2026-10-08 (retrieved 2026-10-10)
- Energy Department Announces New Genesis Mission Awards to Advance Super Intelligence for Science · U.S. Department of Energy · 2026-10-08 (retrieved 2026-10-10)
Revision history (2)
Created Oct 10, 2026. Last reviewed by an editor on Oct 10, 2026. Next scheduled review: Jan 10, 2027.
Cite this page
"Protein Data Bank (PDB)." ContentLora, updated Oct 10, 2026. https://contentlora.com/wiki/protein-data-bank
Spotted an error? Suggest a correction or emailcorrections@contentlora.com.
Keep exploring
- ExplainerHow AI predicts protein structuresWhy a protein's 3D shape matters, how AlphaFold-style models predict it from sequence, and what they still get wrong.
- WikiAlphaFoldAlphaFold is Google DeepMind's AI system for predicting protein and biomolecular structures. Versions, accuracy, database and limits.
- DevelopingAI for science tracker: milestones and what to watchLive tracker of AI for science: protein AI, AI weather models, materials, autonomous labs and AI drug discovery, with dated milestones.
- ExplainerAI for science in 2026: a crash courseA crash course on AI for science: protein prediction and design, AI weather models, materials and drug discovery, autonomous labs.
- ExplainerHow AI designs new proteins, drugs and materialsHow generative AI proposes new proteins, drug molecules and crystals, and why every design still has to be made and tested in a lab.
- WikiCell-free protein synthesisCell-free systems express genes in a tube using extracted cellular machinery. Yields, uses in circuit prototyping and on-demand production, and 2026 funding.