How to Select Compounds After Molecular Docking: From Ranked Results to an Experimental Shortlist

A docking screen ends with a ranked list, but an assay needs a defensible shortlist. Learn how to combine rank, pose quality, chemical diversity, artifact risk, availability, and controls before experimental testing.

A study published online in July 2026 modeled three large-scale docking campaigns in which 2,682 molecules had been synthesized and tested across poor, intermediate, and favorable scoring regions. The authors treated docking score as a noisy correlate of binding affinity and added a term for high-ranking artifacts. Their model reproduced the observed hit-rate curves and predicted that artifacts can increasingly dominate the most extreme scores as libraries grow [1].

That result does not make docking rank useless. In a separate prospective study, Liu and colleagues docked 1.7 billion molecules against AmpC beta-lactamase and tested 1,521. In that campaign, hit rates and affinities improved with docking score, while estimates of hit rate stabilized only after several hundred molecules had been tested [2]. Read together, the studies support a practical position: ranking can contain useful campaign-level signal, but the first rows alone are not a complete experimental-selection strategy.

A docking score can rank a model without selecting an experiment

Separate each signal from the decision it cannot make

Evidence layer Useful question What it does not establish by itself
Docking score and rank Which generated poses scored more favorably under one frozen receptor, search space, engine, and protocol? Measured affinity, activity, selectivity, or freedom from scoring artifacts
Pose review Is the geometry compatible with the binding-site hypothesis and known target chemistry? That the predicted pose is the experimentally bound pose
Chemical-diversity selection Does the assay set cover distinct structural hypotheses rather than repeated analogs? Independent mechanisms or biological activity
Structural and aggregation alerts Which compounds need review or assay-specific controls? That a flagged compound is certainly artifactual—or that an unflagged compound is clean
Availability and assay fit Can the intended chemical entity be obtained and tested under the planned conditions? Lead quality or developability
Experiment What response is measured in the selected assay with its controls? Automatic proof of target mechanism or translation to another assay system

1. Define the experimental decision before opening the ranked list

A shortlist for twelve purchasable compounds is a different decision from a 384-well primary assay or a synthesis campaign. State what the experiment must learn: detect binding, test a functional response, compare chemotypes, challenge a predicted interaction, or establish whether a screening protocol deserves a larger round.

Then freeze the practical boundary. Record the number of assay slots, the target construct and assay format, compound-source requirements, controls, test concentration or dose-response plan, confirmation method, and the result that would justify the next step. There is no universal number of compounds that makes a docking shortlist valid. The right number depends on the decision, assay variability, available controls, and resources.

Minimum experimental brief before compound selection

Decision fieldQuestion to answer
Primary endpointBinding, inhibition, activation, displacement, phenotype, or another defined measurement?
CapacityHow many compounds, concentrations, replicates, and controls can be tested without weakening the design?
Chemical sourceMust compounds be in stock, make-on-demand, already synthesized, or compatible with an internal collection?
Assay constraintsWhich solvent, concentration, detection, stability, and solubility constraints matter?
ConfirmationHow will an initial signal be repeated or tested with an orthogonal method?
Advance ruleWhich measured result, control behavior, and chemical-quality checks permit progression?

2. Preserve the full screening record before filtering

Start selection from a traceable screen, not a detached CSV of successful scores. Keep every submitted ligand identity, preparation outcome, docking failure, receptor record, search box, engine version, scoring function, seed policy, pose, and export connected to the job. Otherwise, filtering can silently change the denominator and make a later shortlist impossible to reconstruct.

The dedicated guide to traceable AutoDock Vina batch docking explains this execution record. If the protocol itself has not yet been challenged, use the pre-screening validation workflow before treating its ranking as selection evidence.

3. Confirm chemical identity, readiness, and availability

A selected row must map back to the exact chemical entity that will be ordered, synthesized, or retrieved. Confirm source ID, structure, stereochemistry, protonation and tautomer policy, salt or fragment handling, prepared 3D state, and PDBQT provenance. Do not let a convenient supplier record silently replace the state that was docked.

Check whether the compound can actually enter the experiment: supplier or synthesis status, expected lead time, quantity, purity information, solvent compatibility, and any known stability or solubility constraint. These checks do not predict biological success; they prevent the computational candidate and tested material from becoming different objects. For file-level problems, use the SDF preflight checklist.

Candidate identity packet

RecordWhy it belongs in the shortlist
Stable source and internal IDsReconnect the result to the submitted and procurable molecule.
Original and prepared structuresExpose transformations made before docking.
Selected chemical stateRecord stereochemistry, protonation, tautomer, charge, and salt decisions.
Preparation and docking terminal statesKeep exclusions and failures visible instead of shrinking the screen.
Supplier or synthesis statusDistinguish a computational proposal from a testable material.
Purity and assay-handling notesGive the experimental team the information needed to plan the test.

4. Use docking rank deliberately—not mechanically

High-ranking candidates deserve attention because a validated campaign may enrich useful chemistry near the top. They should not receive automatic assay priority merely because their numerical order is visible.

Moesgaard and colleagues represented a molecule's position by a log-normalized rank and argued that physically testing compounds across a range of ranks is needed to locate the peak hit rate for a specific campaign [1]. This is not a new AutoDock Vina parameter or a universal instruction to test weak poses. It is an experimental design principle: when resources permit, deliberate rank-band sampling can reveal whether the extreme top is enriched, flat, or contaminated by artifacts. Record the rank-selection rule before assay results are available.

5. Inspect poses and target-specific interactions

Review more than the best numerical mode. Examine severe clashes, buried unsatisfied charges, improbable ligand geometry, exposure of hydrophobic groups, receptor artifacts, and whether the pose supports interactions justified by structural, mutational, or ligand evidence. A generic hydrogen-bond count is not a substitute for a target-specific binding hypothesis.

Prospective studies commonly add this layer. Grotsch and colleagues nominated compounds for synthesis using docking score, predicted pose, chemical novelty, and diversity [3]. Zhou and colleagues filtered and clustered candidates, manually reviewed favorable interactions and geometries, then synthesized and experimentally tested the selected set [4]. These examples support multi-criterion selection in their reported targets; they do not define a universal scoring formula.

Use the separate guide to interpreting AutoDock Vina affinity, RMSD, and poses when the result review is not yet complete.

Pose-review questions to retain

QuestionRecord in the decision
Does the pose occupy the intended site?Site evidence, receptor state, and relevant residues.
Is the geometry physically plausible?Clashes, strained conformations, buried charges, and unresolved concerns.
Are key interactions target-specific?Structural or experimental rationale rather than interaction count alone.
Do alternative modes change the hypothesis?Pose rank, score differences, and the reason one mode was retained.
Do methods disagree?Scoring function, receptor state, or review criterion responsible for the branch.

6. Preserve chemical diversity without turning selection into randomness

If every assay slot is occupied by close analogs, one chemotype can dominate the answer. Define how similarity will be measured before selecting representatives: fingerprint, similarity coefficient, scaffold definition, clustering algorithm, and any threshold. Bemis and Murcko introduced a framework-based way to organize ring and linker cores that remains one useful view of structural diversity [8]. Fingerprint and scaffold views answer different questions, so retain the chosen method rather than reporting that the set was merely “diverse.”

Selecting representatives from distinct clusters broadens the structural hypotheses tested. A small number of close analogs can still be valuable when the explicit goal is early structure-activity information. Keep discovery diversity and series expansion as separate objectives instead of blending them after the results are known.

7. Flag interference and aggregation risk without treating alerts as verdicts

Structural alerts can identify candidates that deserve additional scrutiny. The original PAINS filters were derived from frequent hitters observed in one assay-detection technology, although the authors discussed the patterns across other assays [5]. An alert is therefore a reason to inspect the compound and assay context—not a stand-alone experimental conclusion.

Colloidal aggregation is another source of artifactual inhibition. Irwin and colleagues assembled more than 12,000 known aggregators and built a precedent-based advisor, while explicitly warning that the goal was not to eliminate every potentially aggregating chemotype [6]. Use predictions and similarity warnings to plan appropriate controls with the assay team. A clean result requires experimental behavior, not the absence of a software alert.

Turn a risk signal into a review action

SignalReasonable actionBoundary
PAINS or other substructure alertReview the exact pattern, assay technology, known chemistry, and confirmation plan.A match is not proof of interference.
Similarity to a known aggregatorPlan assay-appropriate aggregation controls and inspect concentration dependence.Prediction is not experimental confirmation.
Reactive or unstable functionalityAssess chemical context, storage, solvent, and assay conditions with a chemist.A generic filter can be overaggressive.
Optical or detection concernCheck the specific readout and use an orthogonal detection method when warranted.Interference depends on the assay.
Solubility concernDefine handling, visible precipitation checks, and a feasible concentration range.A calculated property is not a solubility measurement.

8. Design the shortlist with its controls and confirmation path

A candidate list and an assay plan should be designed together. Include the relevant reference ligand or positive control, suitable negative and technical controls, repeat or dose-response criteria, and the orthogonal measurement required before a signal advances. In the RosettaVS study, initial competition-assay findings for a KLHDC2 candidate were confirmed with biolayer interferometry, and a crystal structure later tested the predicted pose [4]. The sequence of evidence—not one docking score—made the result informative.

Docking and experimental screening can also contribute different chemotypes. In a parallel cruzain campaign, Ferreira and colleagues found partly distinct series from docking and high-throughput screening, while joint evidence helped prioritize well-behaved candidates [7]. The practical implication is not that every project needs both technologies. It is that orthogonal evidence can expose what one method misses.

A practical post-docking shortlist matrix

Assign every assay slot a scientific purpose

Selection role Evidence required Question tested
High-ranking, pose-supported candidate Frozen protocol, favorable rank, plausible reviewed pose, chemical readiness Does the strongest computational evidence survive the assay?
Distinct chemotype representative Declared clustering or scaffold rule plus acceptable pose and risk review Does another structural hypothesis produce a signal?
Alternative rank-band candidate Prespecified rank-sampling rule and the same downstream gates How does experimental hit rate change outside the extreme top?
Method-disagreement candidate Clear disagreement between scoring functions, receptor states, or interaction criteria Which modeling assumption better matches experiment?
Reference or assay control Known identity, expected assay behavior, and documented handling Did the assay and interpretation system behave as intended?

What to hand from docking to the experimental team

Minimum decision-ready handoff

Handoff elementMinimum retained detail
Candidate identitySource ID, structure, selected chemical state, and procurable or synthesized material.
Docking contextReceptor, site, box, engine and version, scoring function, parameters, seed policy, rank, and poses.
Selection rationaleRank band, pose evidence, chemotype or cluster, novelty criterion, and role in the assay set.
Risk reviewStructural alerts, aggregation or assay concerns, solubility notes, and planned controls.
Experimental designAssay, concentrations, replicates, controls, confirmation method, and advancement rule.
Complete denominatorSubmitted, prepared, docked, failed, reviewed, shortlisted, ordered, synthesized, and tested counts.
Outcome linkRaw and processed measurements connected back to the exact candidate and computational record.

The same shortlist decision in two buying contexts

Individual researcher and laboratory requirements

Buyer contextPractical questionEvidence to inspect
Individual researcher or small technical team Can I recover the exact run, inspect poses, explain every selected compound, and export a clean handoff without rebuilding the project? Visible preparation states, settings, poses, failures, history, and exportable results.
Laboratory or organization Can a method owner, chemist, and assay scientist review the same selection record and its boundaries? Versioned inputs, protocol ownership, review annotations, selection criteria, procurement context, and linked experimental outcomes.

Where MolNexus fits—and where the selection remains external

MolNexus 0.1.1 is BioChemIntelli's local Windows desktop application for the protein-small-molecule docking layer. It keeps receptor and ligand preparation review, interaction-box definition, AutoDock Vina execution with Vina or Vinardo scoring, pose inspection, local job history, and exports in one workspace. That supports the upstream evidence record needed before an experimental shortlist is finalized.

MolNexus does not automatically choose experimental candidates, cluster chemical diversity, predict assay interference or ADME, procure compounds, or run an assay. Those decisions remain with the researcher and any chemistry or experimental team. Workflow organization improves visibility and repeatability; it does not make a docking score more accurate.

The current commercial profile is a US$499 one-time license for one Windows PC at a time, perpetual use of the purchased version, and 12 months of updates. MolNexus is coming soon; purchase and download are not open yet.

MolNexus in the post-docking decision

MolNexus supportsRequires an external method or decision
Reviewable receptor and ligand preparation statesFinal chemical identity, procurement, synthesis, and quality control
Explicit search box and Vina or Vinardo settingsTarget-specific protocol validation and assay design
Ranked poses with visual inspectionChemical-diversity clustering, novelty policy, and shortlist allocation
Retained job history and exportsInterference controls, experimental measurements, and biological conclusions

Frequently asked questions

The practical conclusion

The best post-docking shortlist is not the table with the most negative scores. It is a prespecified, traceable allocation of experimental capacity across rank signal, pose plausibility, chemical diversity, artifact risk, availability, and controls.

Preserve every denominator, give each selected compound a scientific purpose, and connect measured outcomes back to the exact computational record. That turns a docking screen into a learning cycle: not proof before the assay, but a stronger way to decide what the assay should test next.

References

  1. Moesgaard L, Shoichet BK, Mailhot O. Modeling the Sensitivity of Large-Scale Virtual Screening to Scoring Function Accuracy, Artifacts, and Library Composition Journal of Chemical Information and Modeling (2026) DOI: 10.1021/acs.jcim.6c01066 Original peer-reviewed modeling study of three large-scale campaigns and 2,682 synthesized and tested molecules. Claims in the article remain descriptive and campaign-specific.
  2. Liu F, Mailhot O, Glenn IS, et al.. The impact of library size and scale of testing on virtual screening Nature Chemical Biology (2025) DOI: 10.1038/s41589-024-01797-w Original prospective study comparing a 1.7-billion-molecule AmpC screen with an earlier 99-million-molecule campaign and testing 1,521 compounds across the score landscape.
  3. Grotsch K, Sadybekov AV, Hiller S, et al.. Virtual Screening of a Chemically Diverse "Superscaffold" Library Enables Ligand Discovery for a Key GPCR Target ACS Chemical Biology (2024) DOI: 10.1021/acschembio.3c00602 Original prospective CB2 receptor study supporting the bounded use of docking score, predicted pose, novelty, and diversity in candidate nomination.
  4. Zhou G, Rusnac DV, Park H, et al.. An artificial intelligence accelerated virtual screening platform for drug discovery Nature Communications (2024) DOI: 10.1038/s41467-024-52061-7 Original prospective study supporting multi-stage filtering, clustering, pose inspection, experimental testing, orthogonal confirmation, and structural follow-up in its reported targets.
  5. Baell JB, Holloway GA. New Substructure Filters for Removal of Pan Assay Interference Compounds (PAINS) from Screening Libraries and for Their Exclusion in Bioassays Journal of Medicinal Chemistry (2010) DOI: 10.1021/jm901137j Original PAINS-filter article. It supports use of structural patterns as review signals while preserving the stated boundary that the filters were derived from one assay technology.
  6. Irwin JJ, Duan D, Torosyan H, et al.. An Aggregation Advisor for Ligand Discovery Journal of Medicinal Chemistry (2015) DOI: 10.1021/acs.jmedchem.5b01105 Original aggregation-advisor study supporting precedent-based aggregation review and the explicit caution against treating the advisor as a universal exclusion rule.
  7. Ferreira RS, Simeonov A, Jadhav A, et al.. Complementarity Between a Docking and a High-Throughput Screen in Discovering New Cruzain Inhibitors Journal of Medicinal Chemistry (2010) DOI: 10.1021/jm100488w Full original prospective comparison supporting the bounded claim that docking and experimental screening contributed partly distinct chemotypes and complementary evidence in one campaign.
  8. Bemis GW, Murcko MA. The Properties of Known Drugs. 1. Molecular Frameworks Journal of Medicinal Chemistry (1996) DOI: 10.1021/jm9602928 Original molecular-framework study supporting a declared scaffold-based view of structural diversity.