How to Choose a Protein Structure for Molecular Docking: Holo, Apo, AlphaFold, or an Ensemble?
Choose a receptor by target state, local pocket evidence, ligand context, and validation. Compare when a holo, apo, AlphaFold, or receptor ensemble best fits a reproducible protein-small-molecule docking campaign.
Choosing a protein structure for molecular docking is not a neutral file-selection step. Its side-chain orientations, pocket volume, missing atoms, bound cofactors, waters, construct boundaries, and conformational state define the physical environment in which every ligand will be sampled and scored. Changing the receptor can therefore change which poses are accessible and which compounds rise in a virtual screen [2-4].
This guide turns that choice into a reviewable decision. It is written for protein-small-molecule docking campaigns, including individual projects and laboratory workflows that must justify why a particular PDB entry, computed model, or receptor ensemble was used.
Start with the biological question, not the download button
What each receptor option can represent
| Receptor option | Most useful when | Main question before use |
|---|---|---|
| Holo experimental structure | A bound ligand stabilizes a relevant pocket and biological state. | Is the ligand, construct, sequence, and state relevant to the compounds and mechanism being studied? |
| Apo experimental structure | The unbound state is biologically relevant or no suitable holo structure exists. | Has the pocket opened, collapsed, or rearranged in a way that affects docking? |
| AlphaFold or another computed model | Experimental coverage is missing, incomplete, or mismatched to the target sequence. | Is the local pocket geometry credible after accounting for confidence, templates, cofactors, and target state? |
| Receptor ensemble | Several conformations represent substantiated states or pocket flexibility that one structure cannot capture. | Does the ensemble add validated information, and how will results across structures be combined? |
1. Lock the target identity and state before comparing structures
Define the object you intend to model before ranking candidate structures. Record the organism, gene or UniProt accession, isoform, residue range, mutations, construct additions, domains, post-translational state, oligomeric assembly, and any required cofactor, metal, partner, or membrane context. For state-dependent targets, also define the relevant active, inactive, open, closed, or allosteric state.
RCSB PDB entry pages expose source organism, mutations, experimental method, assemblies, ligands, and validation data. They also warn that a structure may have multiple biological assemblies, which should be reviewed rather than assuming the asymmetric unit is the intended receptor [1]. A high-resolution structure of the wrong construct or state is still the wrong starting hypothesis.
Minimum target-state specification
| Field | Decision to record |
|---|---|
| Sequence identity | Species, accession, isoform, residue range, mutations, and engineered substitutions. |
| Biological state | Active, inactive, open, closed, substrate-bound, inhibitor-bound, or another defined state. |
| Assembly | Monomer, oligomer, protein complex, or membrane-associated form required by the pocket. |
| Non-protein components | Required metals, cofactors, prosthetic groups, conserved waters, or partner molecules. |
| Binding site | Orthosteric, allosteric, interfacial, covalent, or another explicitly defined site. |
| Screening purpose | Pose hypothesis, focused analog ranking, broad virtual screening, or method development. |
2. Build a candidate inventory before choosing a favorite
Search by a stable target identifier and collect plausible experimental entries before downloading one convenient PDB file. Separate holo and apo candidates, but also annotate the identity of each bound ligand, the experimental method, resolution where applicable, sequence coverage, construct differences, state, cofactors, missing residues, alternate conformations, and release version. Add computed models only after the experimental landscape is visible.
This inventory prevents a common error: comparing one arbitrary crystal structure with one prediction and treating the difference as a verdict on an entire source class. The real decision is among specific receptor hypotheses for one target and one campaign.
3. Judge quality at the binding site, not only across the whole protein
Global resolution, global backbone RMSD, and global model confidence are useful screening signals, but docking is unusually sensitive to local details. Inspect whether binding-site residues are present, whether side chains have supported conformations, whether alternate locations were modeled, whether the ligand fits the experimental density, and whether unresolved loops or terminal segments occlude the search space.
RCSB recommends examining several quality measures together and paying particular attention to mismatches in the region of interest. Its validation reports provide residue-level geometry and experimental-support information; computed structures expose residue-level confidence rather than experimental density [1]. A nominally better global metric does not automatically produce a better docking pocket.
4. When a holo structure is the strongest starting point
A holo structure is often efficient because a ligand has already stabilized an accessible pocket and provides a direct reference for defining the interaction region. In a ten-enzyme comparison, McGovern and Shoichet found that receptor conformation affected virtual-screen enrichment as information moved from holo to apo and modeled structures, although performance remained target-dependent [2].
Relevance matters as much as the holo label. Prefer a structure whose sequence, state, assembly, and binding-site chemistry match the campaign. Check whether the cocrystallized ligand occupies the same subpocket and induces geometry appropriate for the chemical series being considered. A ligand-bound inactive conformation can be an excellent receptor for one hypothesis and a poor receptor for another.
Holo structure: strengths and review points
| Potential advantage | Review before accepting it |
|---|---|
| Preorganized ligand-compatible pocket | Does the bound ligand stabilize the state and subpocket relevant to the new series? |
| Direct site and box reference | Is the ligand well supported experimentally and chemically interpreted correctly? |
| Visible interactions with cofactors or waters | Which components are mechanistically required, displaced, or crystallization artifacts? |
| Experimental coordinates | Are local residues complete, validated, and free from unresolved conflicts near the pocket? |
5. When an apo structure is appropriate
An apo structure may be the only experimental structure, may better represent the biological question, or may expose a state hidden by a particular ligand. Its main risk is not the absence of a ligand by itself; it is that binding-site side chains, loops, or domain positions may not present the geometry needed for the intended compounds.
Compare the apo pocket with relevant holo homologs or target structures where possible. If the pocket differs materially, refinement or an alternative conformation may be justified. One template-guided refinement study reported movement toward holo pockets in 102 of 124 benchmark cases and improved average EF1% from 3.5 for apo receptors to 6.2 for its refined receptors across 40 DUD-E targets [4]. Those values describe that method and those datasets; they are evidence that pocket refinement can matter, not a universal correction factor.
6. When an AlphaFold model expands the campaign
AlphaFold has made structurally informed docking possible for targets and sequences that previously lacked a practical receptor hypothesis. The AlphaFold Protein Structure Database provides residue-level pLDDT values with each model, enabling local confidence review rather than an all-or-nothing judgment [5].
Docking evidence is encouraging but target-specific. In an AutoDock-GPU benchmark of 2,474 PDBbind complexes, Holcomb and colleagues observed higher top-pose success for cocrystal receptors than for AlphaFold2 models, while AlphaFold2 models outperformed the available apo structures in the overlapping set. The authors found the predicted and apo structures complementary and showed that local obstructions, side chains, and cofactors could matter more than global alignment [6].
Prospective evidence is also positive. Lyu and colleagues docked hundreds of millions to billions of molecules against AlphaFold2 and experimental structures for two receptors. The AlphaFold2 campaigns produced hit rates comparable to the experimental structures, and a later cryo-EM structure supported key aspects of one predicted pose [7]. The practical conclusion is not that every prediction replaces every PDB entry. It is that a well-inspected, target-appropriate predicted pocket can be a productive receptor hypothesis when tested with the same discipline as an experimental structure.
For the broader relationship between structure prediction and small-molecule docking, see how AlphaFold expanded the receptor landscape for molecular docking.
Acceptance gate for a computed receptor model
| Check | Evidence to retain |
|---|---|
| Exact sequence and construct | Accession, isoform, residue coverage, mutations, and any modeled or trimmed regions. |
| Local confidence | Residue-level confidence for the pocket, gates, loops, and residues that define interactions. |
| Pocket geometry | Alignment to relevant experimental pockets, side-chain states, volume, clashes, and accessibility. |
| Biological context | Assembly, partner, membrane, cofactors, metals, and state that the single-chain prediction may not encode. |
| Protocol behavior | Redocking, known-ligand recovery, enrichment, or sensitivity results using the intended workflow. |
| Alternatives | Why this model was retained over available apo, holo, homology, or alternative predicted structures. |
7. Use an ensemble only when it answers a defined uncertainty
An ensemble can represent multiple ligand-stabilized pockets, functional states, or substantiated flexibility. It can also multiply preparation work, docking volume, favorable-score opportunities, and interpretation choices. More receptor files do not automatically mean more biological realism.
Rueda, Bottegoni, and Abagyan retrospectively evaluated 1,068 X-ray conformations from 99 proteins. In their benchmark, carefully selected ensembles of three to five experimental conformations improved docking accuracy and binder-decoy separation compared with random selections; adding some conformations was counterproductive [3]. Treat three to five as a result from that benchmark, not a universal ensemble size. Select diversity for a reason and validate the final combination on the target.
What an ensemble protocol must define
| Protocol element | Required decision |
|---|---|
| Inclusion rule | Which structural state or pocket difference earns a conformation a place in the ensemble? |
| Preparation parity | How will protonation, cofactors, waters, missing atoms, and charges be handled consistently? |
| Docking parity | Which box, engine version, scoring function, parameters, and seed policy remain fixed? |
| Aggregation rule | Best score, consensus, state-specific ranking, or another prespecified method? |
| Validation criterion | What target-specific result justifies the added structures and compute? |
| Reporting | How will receptor-specific poses, scores, failures, and provenance remain distinguishable? |
8. Prepare every receptor candidate under a consistent policy
A fair receptor comparison requires more than aligned PDB files. Apply a documented policy for chain and alternate-location selection, missing atoms and residues, protonation, histidine states, metals, cofactors, waters, termini, charges, and PDBQT conversion. If one receptor receives special optimization, record why and do not attribute the resulting difference solely to its source class.
The detailed protein and ligand preparation guide for AutoDock Vina covers this stage. The objective here is comparison integrity: each candidate should differ because of the receptor hypothesis being tested, not because its preprocessing was undocumented.
Receptor preparation record
| Record | Minimum detail |
|---|---|
| Source | PDB or model identifier, version, download date, and original URL. |
| Selection | Chosen assembly, chains, alternate locations, residues, and retained non-protein components. |
| Repairs | Added atoms or residues, modeled gaps, resolved clashes, and any structural refinement. |
| Chemistry | pH assumption, protonation choices, histidine states, metals, cofactors, waters, and charges. |
| Conversion | Software, version, parameters, warnings, exclusions, and final receptor checksum. |
9. Run a receptor-selection pilot before screening the library
Use a small panel of known ligands, suitable decoys or comparators, and any cognate ligand with a known pose. Keep the ligand preparation policy, search region, engine version, scoring function, exhaustiveness, and evaluation metrics fixed across receptor candidates. Compare pose recovery where a reference pose exists, known-ligand ranking or enrichment where the data support it, failure patterns, and sensitivity to reasonable box or preparation choices.
A protocol article from Bender and colleagues likewise treats receptor preparation and calibration with known ligands as central stages of large-scale docking [8]. Use the molecular docking validation workflow to define the pilot before observing which receptor wins.
A practical receptor-selection hierarchy
- Define the exact target state. Exclude structures that cannot represent it, regardless of headline quality.
- Start with a relevant, locally validated holo structure when one exists. It is usually the most direct pocket hypothesis.
- Retain credible alternatives. An apo or predicted model may better match the sequence, state, construct, or pocket question.
- Compare local pocket evidence. Inspect residues, density or confidence, ligand quality, cofactors, waters, and unresolved regions.
- Prepare candidates consistently. Keep every transformation and exception traceable.
- Run a prespecified pilot. Select the simplest receptor or small ensemble that performs adequately for the intended use.
- Record the rejected options. A reproducible choice includes why plausible alternatives were not used.
Receptor selection register
| Candidate | Decision | Evidence retained |
|---|---|---|
| Primary receptor | Chosen for the production workflow. | Target-state match, pocket review, preparation record, and pilot result. |
| Alternative receptor | Retained for sensitivity analysis or a distinct biological state. | Specific uncertainty it represents and separate receptor-level results. |
| Excluded receptor | Not used in production docking. | Sequence, state, quality, pocket, component, or validation reason for exclusion. |
What individual researchers and organizations should require
The same decision in B2C and B2B workflows
| Buyer context | Practical requirement | Main fit question | Evidence and current next step |
|---|---|---|---|
| Individual researcher paying for personal use | Compare plausible receptor inputs without losing the exact preparation, box, settings, poses, or result history. | Does a local Windows workflow and a one-PC license fit how I work? | Inspect input provenance, preparation states, run settings, receptor-specific outputs, exports, and the published price and license terms. The product page can record a launch notification. |
| Laboratory, core facility, or company controlling the purchase | Let different scientists review and reproduce why one receptor state entered a campaign. | Does the currently documented single-PC offer satisfy deployment and procurement needs? | Inspect the versioned receptor register, controlled protocol, validation criteria, review notes, retained alternatives, and the current product specifications. Team, site, and institution-wide terms are not published. |
Where MolNexus fits - and where receptor selection remains scientific judgment
MolNexus 0.1.1 is BioChemIntelli's local Windows desktop application for protein-small-molecule docking. It brings receptor and ligand preparation review, interaction-box definition, AutoDock Vina 1.2.7 execution with Vina or Vinardo scoring, pose inspection, local SQLite job history, and exports into one workspace. Researchers can use separate controlled jobs to compare receptor candidates while retaining the execution context.
MolNexus does not determine the biologically correct receptor, generate or optimize conformational ensembles, automatically rank PDB entries, or prove that one structure is experimentally valid. Those decisions require target knowledge, structural inspection, and protocol validation. MolNexus organizes the docking layer after the receptor hypothesis is defined; it does not replace that hypothesis.
The current commercial profile is a US$499 one-time license for one Windows PC at a time, perpetual use of the purchased version, and 12 months of updates. MolNexus is coming soon; purchase and download are not open yet.
The announced offer is a single-PC license. BioChemIntelli has not published team, site, floating, or institution-wide licensing or procurement terms, so this article does not present them as available.
MolNexus in a receptor-comparison workflow
| MolNexus supports | Remains an external scientific decision |
|---|---|
| Reviewable receptor import and preparation state | Target identity, biological state, assembly, and source selection |
| Explicit interaction box and Vina or Vinardo settings | Target-specific validation design and receptor acceptance criteria |
| Pose inspection for each controlled job | Structural refinement, ensemble generation, and cross-receptor aggregation policy |
| Local job history and result exports | Experimental interpretation and biological conclusions |
Frequently asked questions
The practical conclusion
The best receptor is not simply the newest prediction, the highest-resolution structure, or the first ligand-bound entry. It is the simplest well-supported structural hypothesis that matches the target state, survives local pocket review, and behaves adequately under a prespecified validation protocol.
Start from a relevant holo structure when it offers the clearest pocket evidence. Use apo and AlphaFold models as valuable, sometimes complementary options when they better address the sequence or state question. Add an ensemble only when its extra conformations earn their place. Above all, preserve the receptor decision, preparation, settings, and results as one traceable record.
References
- RCSB Protein Data Bank. Assessing the Quality of 3D Structures RCSB PDB Documentation (2026) Current authoritative RCSB guidance on experimental validation, residue-level quality, ligand quality, and computed-model confidence. Last updated in 2026.
- McGovern SL, Shoichet BK. Information Decay in Molecular Docking Screens against Holo, Apo, and Modeled Conformations of Enzymes Journal of Medicinal Chemistry (2003) DOI: 10.1021/jm0300330 Full original paper hosted by the authors laboratory. It compares database docking against holo, apo, and modeled structures for ten enzyme binding sites.
- Rueda M, Bottegoni G, Abagyan R. Recipes for the Selection of Experimental Protein Conformations for Virtual Screening Journal of Chemical Information and Modeling (2010) DOI: 10.1021/ci9003943 Full original study of 1,068 X-ray conformations from 99 proteins and target-specific single-receptor and ensemble performance.
- Guterres H, Park SJ, Jiang W, Im W. Ligand Binding Site Refinement to Generate Reliable Holo Protein Structure Conformations from Apo Structures Journal of Chemical Information and Modeling (2021) DOI: 10.1021/acs.jcim.0c01354 Full original benchmark of a template-guided MD refinement method on DUD-E and Gunasekaran apo-holo datasets. Numerical claims are kept specific to that method and those datasets.
- AlphaFold Protein Structure Database. Frequently Asked Questions: Model outputs and confidence EMBL-EBI and Google DeepMind (2026) Official documentation for AlphaFold DB model outputs and residue-level pLDDT confidence interpretation.
- Holcomb M, Chang YT, Goodsell DS, Forli S. Evaluation of AlphaFold2 Structures as Docking Targets Protein Science (2023) DOI: 10.1002/pro.4530 Full original AutoDock-GPU study comparing cocrystal, AlphaFold2, and available apo structures across a large PDBbind-derived set.
- Lyu J, Kapolka N, Gumpper R, et al.. AlphaFold2 Structures Guide Prospective Ligand Discovery Science (2024) DOI: 10.1126/science.adn6354 Full original prospective comparison of large-library docking against AlphaFold2 and experimental structures for two receptors, with synthesis, testing, and structural follow-up.
- Bender BJ, Gahbauer S, Luttens A, et al.. A Practical Guide to Large-Scale Docking Nature Protocols (2021) DOI: 10.1038/s41596-021-00597-z Peer-reviewed practical protocol supporting receptor preparation, calibration, controls, and documentation before large-scale docking.