How to Interpret the Confidence of Predicted Structures in AlphaFold
Why This Matters
AlphaFold has significantly broadened the accessibility of structural biology by predicting possible three-dimensional protein structures from amino-acid sequences. Proteins lacking experimental structures can be analyzed for domain and residue arrangements to formulate hypotheses regarding mutations or experiments. However, interpreting predicted coordinates as a single, fixed structure within the actual cell causes the model to assert answers to questions it did not address.
AlphaFold2, AlphaFold-Multimer, and AlphaFold3 differ in their targets and learning-output contracts. To compare results, one must record the specific model and database used, as well as the template and MSA settings and release version.
What Predictions Generate
The model learns spatial constraints between residues from sequence and evolutionary information to generate coordinates. Rather than examining only the output PDB structure, interpret the confidence metrics alongside it. The prediction represents the arrangement deemed most plausible by the model, rather than experimental observations such as electron density or NMR restraints.
The same sequence may adopt different conformations depending on ligands, binding partners, membrane environments, pH, and post-translational modifications. Intrinsically disordered regions may be difficult to represent as a single stable structure.
Reading pLDDT
pLDDT summarizes the model confidence in the local geometry around each residue. High values indicate that the backbone and side-chain placements are relatively reliable, while low values may indicate regions of uncertainty or potential disorder. Low pLDDT should not be taken as definitive evidence of experimental failure or non-functional regions.
High pLDDT is also not evidence that the prediction performs a biological function. Even if the local fold is correct, the overall domain orientation or complex state may be incorrect.
PAE and Domain Arrangement
PAE (predicted aligned error) displays the expected error in the position of other residues when aligned based on a single residue as a matrix. If the intra-domain PAE is low but the inter-domain PAE is high, it can be interpreted that the fold of each domain is reliable while their relative orientation is flexible or uncertain.
PAE is particularly useful for avoiding interpretation of proteins connected by long linkers into a fixed domain arrangement based on a single PDB snapshot. Multimer interfaces are also evaluated in conjunction with interface confidence and alternative stoichiometry.
Example of a Small Variant
Even if a patientโs missense variant is located at a highly confident core conserved residue and appears to disrupt local packing, pathogenicity is not definitively established. Confirmation requires data on protein expression and stability, enzymatic activity, cellular phenotypes, and clinical segregation. Conversely, variants in the low-confidence tail region cannot be dismissed, as they may alter linear motifs or regulatory modifications.
Structural information serves as a guide for designing functional assays but does not constitute the functional assay results themselves.
How to Use Predictions Effectively
Narrow the question to a scope that the model can support, such as domain boundary, candidate interface, or mutagenesis site. Compare homolog structures and experimental data with multiple predictions and confidence scores. Do not assert binding affinity or drug efficacy based solely on docking scores or visual clashes.
Common Misconceptions
- High confidence does not equate to high biological importance.
- Low confidence does not necessarily indicate an incorrect sequence or lack of function.
- Structural similarity does not automatically guarantee identical substrate or function.
- Predicted complexes do not confirm in-cell binding, concentration, or stoichiometry.
- Output differences between model versions do not reflect changes in experimental structures.
Interpretive Boundaries
AlphaFold results aid in hypothesis generation and experimental design. Cellular environment, dynamics, alternative states, binding affinity, and clinical pathogenicity must be cross-validated with biochemical, cellular, and structural evidence.
Reading in Context
The relationship between prediction and functional evidence is linked to functional assay and variant interpretation, and the clinical classification of variants is connected to the Genotypeโphenotype evidence concept.