AlphaFold Database Now Includes Protein Pair Interactions
AlphaFold free access protein structure database, widely used for molecular biology research & reportedly containing almost every currently known protein, has been expanded to include predicted protein–protein interactions in the form of homodimers—a structure formed when two identical protein molecules bind together and function as a single unit.
These matching proteins attach through specific interactions and often become more stable or active in this paired form, which is required for many biological processes. For example, enzymes such as HIV-1 protease only become active when two copies of the same protein assemble into a dimer.

AF-0000000066503175: Homodimer of Transcription elongation factor Eaf N-terminal domain-containing protein. Credit: AlphaFold Protein Structure Database
The AlphaFold Protein Structure Database is a repository of roughly 200 million individual precomputed protein predictions that were generated by running AlphaFold models at scale on large protein-sequence collections.
The AlphaFold models family are AI systems that take molecular inputs (most commonly amino-acid sequences) and output predicted 3D coordinates plus confidence metrics. AlphaFold 3 extends the scope to complexes that can include proteins, nucleic acids, small molecules, ions, and modified residues.
Since its release in 2021, the AlphaFold database has until now focused on individual protein structures. These monomer predictions capture the 3D shape of single proteins, but do not account for how proteins interact with each other—a key determinant of biological function.
Protein complex prediction introduces additional computational complexity, a layer of functional context that was previously missing from the resource. Initial computational runs generated around 30 million possible homodimer structures, which were then filtered down to 1.7 million entries based on confidence thresholds.
The expansion was carried out by a consortium including EMBL’s European Bioinformatics Institute, Google DeepMind, Seoul National University, and NVIDIA. The group focused on proteins from 20 well-studied organisms, including humans, mice, yeast, and several disease-associated bacteria, including the World Health Organization’s priority pathogens list.
Further updates are expected to include heterodimers, involving interactions between different proteins. The consortium has already generated around eight million such predictions, derived from known interaction datasets, although only a subset will be incorporated into the main AlphaFold database based on quality criteria.
Despite filtering, variability in prediction accuracy remains. Some predicted complexes may not correspond to biologically relevant interactions, particularly in cases where binding is transient or context-dependent. Experimental validation remains necessary for confirming functional interactions in downstream applications such as drug discovery or pathway analysis.
Bridging Prediction and Design
While the AlphaFold database update adds a new layer of biological context, it still remains primarily a structural reference resource. Isomorphic Labs, which built the AlphaFold 3 AI engine together with Google DeepMind, is moving beyond simply predicting what proteins and protein complexes look like.
IsoDDE, which was introduced in February 2026, predicts how drugs can bind to proteins and helps design and prioritize new therapeutic molecules based on their likely effectiveness.
For readers who want a broader map of how AI is starting to move from structure prediction into design, our deep dive Protein Language Models: Builders & Pharma Deals breaks down how protein language models learn from amino-acid sequences, which teams are building the leading systems, and where pharma is already partnering or licensing capabilities.
It also flags the main technical constraints that still limit practical deployment, including the sequence–structure gap, protein dynamics, and the difficulty of connecting sequence-only models to binding and function, which is the same boundary newer systems like IsoDDE are starting to push on.
Topic: AI in Bio