AlphaFold Didn’t Replace Molecular Docking—It Unlocked Its Next Era
AlphaFold has dramatically reduced the structural bottleneck that kept many protein targets outside docking studies. Learn how predicted structures become receptor starting points for protein–small-molecule docking and virtual screening with AutoDock Vina.
For decades, many protein–small-molecule docking projects stopped at their first practical question: is there a usable receptor structure? Experimental structure determination remains indispensable, but it can require substantial time, specialized facilities, and a construct that behaves well experimentally. Promising proteins without an adequate structure were therefore difficult to investigate with structure-based virtual screening.
AlphaFold changed that starting condition. It made plausible protein models available at proteome scale and transformed receptor scarcity into receptor assessment. This article focuses specifically on the resulting workflow for protein–small-molecule docking: AlphaFold supplies a structural hypothesis for the protein receptor; AutoDock Vina evaluates how new small-molecule ligands can fit within a defined site. They solve consecutive parts of the same study rather than replacing one another.
From structural scarcity to a much larger target universe
Known protein sequences grew far faster than experimentally determined structures. AlphaFold2 narrowed that gap with strong CASP14 performance and a calibrated per-residue confidence estimate [2]. The database-scale effect matters just as much as that accuracy milestone: the 2024 AlphaFold Database release reported structural coverage for more than 214 million protein sequences [1].
A researcher can now begin structural assessment for many understudied proteins, pathogen targets, isoforms, and non-model organisms without first waiting for a deposited experimental structure. The practical question has shifted from “is any receptor model available?” to “which available model and biological state provide the best starting point for this docking question?” That shift substantially expands the target universe accessible to protein–small-molecule studies.
How AlphaFold and AutoDock Vina divide the workflow
| Workflow stage | AlphaFold contribution | AutoDock Vina and docking contribution |
|---|---|---|
| Target access | Provides plausible protein coordinates when no suitable experimental structure is available. | Uses an existing three-dimensional receptor; it does not predict the protein fold. |
| Receptor assessment | Provides pLDDT and PAE estimates for evaluating local and inter-domain confidence. | Depends on a receptor whose pocket geometry, chemistry, and biological state have already been assessed. |
| New small-molecule ligands | Canonical AlphaFold2 Database models supply the protein structure, not a ligand-screening workflow. | Accepts prepared new ligands and explores their poses inside a defined search space. |
| Search and comparison | Expands the number of proteins for which structural questions can be formulated. | Samples ligand poses and produces docking scores under a documented protocol. |
| Scientific outcome | A receptor hypothesis with explicit structural confidence. | A ranked set of testable protein–ligand pose hypotheses. |
Can AlphaFold structures be used for protein–small-molecule docking?
The decision is local. Docking is sensitive to the geometry of a small region: cavity shape, rotamers of interacting residues, charge state, and sometimes the relative placement of domains. AlphaFold DB reports pLDDT as a per-residue estimate of local confidence and PAE as an estimate of confidence in the relative positions of residue pairs or domains [1,4]. Inspecting those metrics around the proposed site is more informative than quoting one protein-wide average.
Confidence is also not biological-state evidence. A model can be locally well formed while representing the wrong functional state for a particular ligand. Proteins may switch between open and closed forms, reorganize loops on binding, assemble with partners, or depend on a membrane, ion, cofactor, or post-translational modification. Those questions require external evidence.
Why an AlphaFold structure still needs receptor preparation
A canonical AlphaFold DB AlphaFold2 coordinate file contains protein atoms, but receptor preparation is a scientific interpretation step. Important limitations include:
- Conformational state: the prediction may not represent the ligand-compatible or experimentally relevant state.
- Pocket side chains: small rotamer differences can change hydrogen bonds, steric fit, and docking rank.
- Non-protein chemistry: canonical AlphaFold2 database models do not place cofactors, metals, ligands, ions, nucleic acids, or post-translational modifications [4].
- Flexible or disordered regions: low-confidence coordinates are still present in downloaded files and should not be interpreted as equally reliable.
- Biological assembly: a monomer may expose a cavity that is occluded, reshaped, or created by oligomerization.
- Docking representation: protonation, hydrogen placement, charges, atom types, water decisions, and PDBQT conversion remain downstream choices.
Deleting every heteroatom from a homologous structure or accepting every predicted atom without review are both mechanical shortcuts. The right preparation depends on the target mechanism and on the evidence available for the site.
From AlphaFold model to Vina-ready receptor
1. Verify identity and biological state
Confirm the UniProt accession, sequence version, isoform, organism, construct boundaries, signal-peptide or propeptide processing, and any engineered mutations. Then ask which state the experiment or hypothesis requires: monomer or complex, active or inactive, apo or ligand-bound, soluble or membrane-associated.
Save the downloaded model version and a checksum of the coordinate file. If the input cannot be reconstructed, later parameter comparisons will not be genuinely reproducible.
2. Inspect confidence where docking will occur
Map pLDDT onto the structure and list the residues that define the proposed pocket. AlphaFold DB describes values above 90 as its highest-confidence local category, 70–90 as generally reliable backbone predictions, 50–70 as low confidence, and values below 50 as frequently associated with disordered or unstructured regions [1,4]. These are interpretation categories, not a universal docking cutoff.
Review PAE when multiple domains or modeled chains jointly form the site. Low local uncertainty within each domain does not guarantee that their relative orientation is reliable. If the pocket depends on uncertain packing, compare alternative structures or an ensemble and seek experimental restraints rather than treating one receptor conformation as definitive.
3. Compare experimental and homologous structures
Search the Protein Data Bank for the same protein, domains, close homologues, bound ligands, conserved catalytic residues, and known conformational states. A global structural alignment is only the start. Examine pocket volume, side-chain rotamers, loop placement, conserved waters, metals, and ligand-induced changes.
When an appropriate ligand-bound experimental structure exists, it will often be the more informative docking template. The AlphaFold model may still add value as an alternative state or for missing regions, but the choice should be justified rather than decided by novelty.
4. Repair chemistry without erasing uncertainty
Resolve missing or ambiguous side chains, clashes, alternate residue identities, disulfides, termini, and protonation states with a documented method. Add essential metals or cofactors only when biological and structural evidence supports their presence and geometry. Decide which waters to retain based on their role, not through an automatic “remove all waters” rule.
Energy minimization can remove local strain, but aggressive refinement can create false confidence. Preserve the original model, record every modification, and compare the prepared receptor with its source. A polished structure is not necessarily a more accurate structure.
5. Define the search space from evidence
Center and size the docking box using a co-crystallized ligand, mutagenesis, catalytic residues, conserved motifs, pocket-detection evidence, or a clearly stated hypothesis. The box must cover the intended site without becoming so large that sampling is diluted across irrelevant surface.
If the site itself is unknown, blind docking may generate hypotheses, but it is a different and more uncertain question than screening a defined pocket. Report it accordingly. Do not present a visually plausible surface pose as proof that the site is functional.
A reproducible AlphaFold-to-AutoDock Vina workflow for small molecules
Minimum workflow and decision record
| Stage | Record | Decision gate |
|---|---|---|
| 1. Provenance | Accession, sequence and model version, source URL, download date, checksum. | Does this model represent the intended target construct? |
| 2. Confidence review | Pocket-residue pLDDT, relevant PAE regions, uncertain loops or interfaces. | Is the proposed site geometrically interpretable? |
| 3. Structural cross-check | Experimental structures, homologues, states, ligands, conserved residues. | Is this receptor state defensible for the question? |
| 4. Receptor preparation | Repairs, hydrogens, protonation, charges, cofactors, metals, waters, output hash. | Is the chemistry internally consistent and documented? |
| 5. Ligand preparation | Source, stereochemistry, tautomer, protomer, charge method, conformers. | Do inputs represent plausible chemical states? |
| 6. Search space | Grid center, dimensions, spacing where applicable, and biological rationale. | Does the box test the stated site hypothesis? |
| 7. Calibration | Redocking where possible, known actives, inactives or decoys, parameter sensitivity. | Can the protocol recover useful signal for this target? |
| 8. Production run | Vina version, seed, exhaustiveness, modes, energy range, hardware, logs. | Can another researcher repeat the same calculation? |
| 9. Pose review | Interactions, clashes, strain, conserved contacts, alternative poses and ranks. | Is the pose chemically plausible beyond its score? |
| 10. Validation | Orthogonal computation, prospective assay, controls, acceptance criteria. | What evidence could falsify the docking hypothesis? |
AlphaFold is upstream of docking—not a replacement for it
The workflow described here uses canonical AlphaFold2 Database models as potential receptor sources. Those entries provide protein coordinates and confidence information; they do not accept a library of new ligands, define a docking box, search ligand poses, or return Vina-style docking scores. Those are downstream docking operations.
AlphaFold 3 can jointly predict selected biomolecular complexes, including protein–small-molecule complexes [5]. That is a different task from screening an enumerated ligand library against a prepared receptor. For the opportunity discussed in this article, the division remains clear: AlphaFold broadens structural access to protein targets, while a docking engine evaluates new ligands systematically.
Can AlphaFold-derived receptors support virtual screening?
Published results show both the scale of the opportunity and the importance of the downstream docking protocol. In one AlphaFold2-enabled antibiotic benchmark, AutoDock Vina predictions across 12 experimentally tested targets produced a mean auROC of 0.48. Selected machine-learning rescoring functions raised the mean as high as 0.63. For the available target subset, docking against experimentally determined structures produced a similarly weak mean auROC of 0.46 [6]. Receptor availability opened the analysis, but docking and scoring quality still determined whether the workflow separated actives from inactives.
Prospective evidence demonstrates what becomes possible when a predicted receptor supports a well-designed campaign. A 2024 Science study used DOCK3.8 to screen the same 490 million molecules against AlphaFold2 and crystal structures of the σ2 receptor, and more than 1.6 billion molecules against AlphaFold2 and cryo-EM structures of the 5-HT2A receptor. Reported hit rates were 54% versus 51% for σ2 and 26% versus 23% for 5-HT2A, comparing the AlphaFold2 and experimental-structure campaigns, respectively [7]. For these two targets, predicted structures enabled ligand-discovery campaigns at a scale that would previously have depended on experimental receptor availability.
Checklist before using an AlphaFold model as a Vina receptor
- The protein accession, sequence, isoform, construct, and biological state are explicit.
- The model source, version, download date, and file checksum are preserved.
- pLDDT has been inspected for the pocket residues, not only averaged across the protein.
- PAE or interface confidence supports any domain or chain arrangement that shapes the site.
- Relevant experimental structures, homologues, motifs, and bound ligands have been compared.
- Every repair, protonation choice, cofactor, metal, water, and charge method is documented.
- Ligand stereochemistry, tautomeric and protonation states, and preparation settings are recorded.
- The docking box has coordinates, dimensions, and an evidence-based rationale.
- Vina version, random seed, exhaustiveness, number of modes, and output logs are preserved [8].
- Known ligands, inactives, decoys, redocking, or another target-relevant calibration has been attempted where feasible.
- Top poses have been inspected for clashes, strain, unsupported interactions, and ranking instability.
- The claim is limited to a testable hypothesis and is paired with an orthogonal or experimental validation plan.
Frequently asked questions
Why AlphaFold creates a new era for protein–small-molecule docking
AlphaFold’s most important contribution to docking is the transition from structural scarcity to target selection. Researchers can now assess many proteins that were previously outside structure-based studies, compare receptor families at scale, and move promising targets into small-molecule docking workflows much earlier.
The new bottleneck is no longer always obtaining any protein structure. It is selecting the most relevant model and state, preparing the receptor and ligands consistently, defining an evidence-based binding site, and preserving every decision from input to ranked pose. That is a far more productive problem—and one that reproducible docking software can help researchers solve.
References
- Varadi M, Bertoni D, Magana P, et al. AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences Nucleic Acids Research 52(D1):D368–D375 (2024) DOI: 10.1093/nar/gkad1011 Primary database paper for coverage, confidence metrics, data access, and documented exclusions.
- Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold Nature 596:583–589 (2021) DOI: 10.1038/s41586-021-03819-2
- Varadi M, Anyango S, Deshpande M, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models Nucleic Acids Research 50(D1):D439–D444 (2022) DOI: 10.1093/nar/gkab1061
- EMBL-EBI. AlphaFold Protein Structure Database: Frequently Asked Questions AlphaFold Protein Structure Database Official guidance on pLDDT, PAE, model contents, and limitations; accessed 24 July 2026.
- Abramson J, Adler J, Dunger J, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3 Nature 630:493–500 (2024) DOI: 10.1038/s41586-024-07487-w
- Wong F, Krishnan A, Zheng EJ, et al. Benchmarking AlphaFold-enabled molecular docking predictions for antibiotic discovery Molecular Systems Biology 18:e11081 (2022) DOI: 10.15252/msb.202211081
- Lyu J, Kapolka N, Gumpper R, et al. AlphaFold2 structures guide prospective ligand discovery Science 384(6702):eadn6354 (2024) DOI: 10.1126/science.adn6354
- Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings Journal of Chemical Information and Modeling 61(8):3891–3898 (2021) DOI: 10.1021/acs.jcim.1c00203