AI-Generated Ligands Need More Than a Score: What to Evaluate in a Molecular Docking Workflow
Molecular docking for AI-generated ligands needs chemical preparation, controlled protocols, pose review, and traceable results. See a documented CDK2 case in MolNexus.
This guide is for computational chemists, molecular modelers, and research teams receiving more generated molecules than they can evaluate manually. The immediate problem is no longer proposal volume. It is turning those proposals into a defensible screening record without allowing invalid inputs, inconsistent preparation, or detached score tables to obscure what actually happened.
Molecular docking for AI-generated ligands should therefore be evaluated as a sequence of gates: source provenance, chemical identity, ligand preparation, receptor and pocket definition, controlled docking, pose review, and traceable export. Software earns its place in that sequence by making these transitions explicit.
Why an AI-generated molecule is not yet a docking candidate
A generator may output a molecular graph, SMILES string, three- dimensional coordinates, or a structure already positioned near a pocket. None of those representations automatically answers every downstream question. A candidate can still require checks for valence, bond orders, aromaticity, stereochemistry, protonation, tautomeric state, duplicates, disconnected fragments, and compatibility with the selected preparation tool.
Drug-likeness and synthetic-accessibility filters answer different questions again. Published generative workflows commonly combine generation with chemoinformatic filtering, docking, more intensive simulations, and experimental follow-up rather than treating one model output as a finished candidate [1,5,6]. The workflow is layered because each stage removes a different kind of uncertainty.
What each evaluation stage can—and cannot—answer
| Stage | What it can contribute | What it does not establish by itself |
|---|---|---|
| Generative model | Propose novel molecular representations under the model's training and conditioning scheme. | Correct chemistry in every downstream tool, target binding, synthesis, or biological activity. |
| Chemical checks and preparation | Test whether a representation can be interpreted, assign a selected state, and create docking-ready input. | That the chosen state is the only relevant state or that the molecule will bind. |
| Molecular docking | Generate candidate poses and scores within one receptor, search space, engine, and parameter set. | Measured binding affinity, selectivity, kinetics, cellular activity, or synthetic feasibility. |
| Post-docking analysis | Inspect interactions, compare poses, apply controls, and prioritize candidates for deeper evaluation. | A biological conclusion without the additional evidence required by the study. |
| Experimental evaluation | Measure the selected physical or biological endpoint under a defined assay. | That every computational assumption was correct or that one assay answers every development question. |
A seven-gate workflow for AI-generated ligands
1. Preserve generation provenance before changing the molecule
Assign a stable internal identifier and retain the unmodified source representation. Record the generator or dataset, version, conditioning target, generation run, source filename, and any supplied score or property. If the molecule came from a published dataset, retain the DOI, version, license, and file checksum.
This boundary matters because preparation may change coordinates, hydrogen counts, bond perception, charge representation, or file format. A later reviewer should be able to distinguish what the generator proposed from what the docking workflow prepared.
2. Apply an explicit chemical-identity gate
Check whether the molecule can be parsed and sanitized under the tools that will actually prepare and dock it. Review valence, aromaticity, disconnected components, stereochemical specification, formal charge, and duplicate identity. Decide how salts, mixtures, alternate stereoisomers, tautomers, and protonation states will be handled.
A failure at this stage is useful evidence, not an inconvenience to hide. Record the source molecule, terminal state, and reason. Correcting a representation may be scientifically reasonable, but the correction should create a new documented representation rather than overwrite the original proposal.
3. Prepare every accepted ligand under one declared protocol
Docking engines require more than a molecule name. Preparation establishes the representation used by the engine: three-dimensional geometry, hydrogen treatment, atom typing, charges, and rotatable-bond definitions. Apply the same declared policy to the batch unless a molecule-specific exception is justified and recorded.
Preserve both the source and prepared file. For batch work, report the number selected, prepared, rejected, and intentionally excluded. The denominator in a docking summary should be the number that actually reached the engine, not the number originally generated.
4. Define the receptor, pocket, and search space independently
A target label such as CDK2 does not uniquely specify a docking experiment. Record the exact structure source, chain, conformation, retained components, missing or altered residues, preparation decisions, and prepared receptor file. Then record the search-box center, size, and structural rationale.
The pocket used to condition a generative model may provide useful context, but it does not remove the need to define the receptor and search space used by the docking engine. If a predicted receptor is under consideration, see the AlphaFold-to-AutoDock Vina workflow guide for the upstream structural checks.
5. Run a controlled docking protocol
Record the docking engine and version, scoring function, box coordinates, exhaustiveness, number of modes, energy range, CPU setting, seed behavior, and every parameter changed from the chosen protocol. AutoDock Vina supports batch docking of multiple ligands, but a consistent configuration and prepared inputs remain essential for interpreting the resulting rows [4,7].
Use controls appropriate to the target and decision. Depending on the study, these may include redocking, known actives and decoys, alternate receptor states, repeated seeds, or target-specific enrichment analysis. A production workflow should make adding those controls possible rather than presenting an uncalibrated score threshold as universal.
6. Review poses in receptor context
The first ranked score is a starting point for inspection, not a substitute for it. Review whether the pose occupies the intended site, whether important interactions are geometrically plausible, whether strained or exposed features deserve attention, and whether alternative modes tell a materially different story.
Keep each pose connected to the source ligand, prepared ligand, receptor state, search box, and run configuration. The guide to organizing receptors, ligands, and poses describes that record structure in detail.
7. Preserve history and export a decision-ready record
A useful result table includes stable ligand identity, preparation status, docking status, top score, pose count, and links to the underlying outputs. Preserve failures alongside successes. Add the protocol, software versions, receptor record, and review notes before exporting a shortlist.
This is where workflow software can create commercial value: fewer manual handoffs, clearer terminal states, faster recovery of prior work, and a result package that can be audited without reconstructing the run from loosely named files.
Real workflow: eight YuelDesign ligands against apo CDK2
The source was version v2 of the public YuelDesign dataset deposited
on Zenodo under CC BY 4.0 [1,2]. Its
cdk2_yueldesign.tar archive contains 210 generated CDK2
structures in SDF form. Before any preparation result was known, we
selected the *_1.sdf entry for each even deposited size
label from 18 through 32.
For the receptor, we downloaded the 1.34 Å X-ray structure of apo human CDK2, PDB 4EK3 [3]. A Kabsch rigid-body fit aligned the complete 4EK3 structure to 251 shared atoms in the deposited YuelDesign pocket coordinate frame, yielding 0.000498 Å RMSD. This transformed coordinates, not receptor conformation. MolNexus then prepared the receptor and recorded that incomplete residues A:36, A:73, A:74, and A:75 could not be parameterized and were excluded.
Prespecified ligand outcomes in the documented workflow
| Deposited source | Preparation | Docking | Best Vina score (kcal/mol) | Retained poses |
|---|---|---|---|---|
size18_1 |
Prepared | Completed | −6.536 | 5 |
size20_1 |
Rejected: implicit-hydrogen representation was not accepted by the preparation path | Not run | — | 0 |
size22_1 |
Prepared | Completed | −6.438 | 4 |
size24_1 |
Prepared | Completed | −6.767 | 5 |
size26_1 |
Prepared | Completed | −6.991 | 5 |
size28_1 |
Prepared | Completed | −6.916 | 5 |
size30_1 |
Rejected: aromatic bond assignment could not be resolved by the preparation path | Not run | — | 0 |
size32_1 |
Prepared | Completed | −7.006 | 5 |
The size22_1 run returned five raw modes.
One mode was excluded by the workflow's nominal-box check because
one heavy atom extended 0.119 Å beyond the declared search box,
leaving four retained poses.
Documented protocol behind the interface capture
| Source dataset | YuelDesign CDK2 molecules, Zenodo 17702010 v2, CC BY 4.0 |
|---|---|
| Selection rule | Before preparation, select *_1.sdf for even deposited size labels 18–32; 8 of 210 candidates |
| Receptor | Apo human CDK2, PDB 4EK3, X-ray resolution 1.34 Å |
| Coordinate alignment | Kabsch fit of full 4EK3 to 251 shared pocket atoms; 0.000498 Å RMSD |
| Search box | Center (0.0, 28.7, 9.5) Å; size 24 × 20 × 20 Å |
| Engine | AutoDock Vina 1.2.7, Vina scoring function |
| Run settings | Exhaustiveness 16; 5 modes; 3 kcal/mol energy range; seed 20260725; 1 CPU |
| Product workflow | MolNexus 0.1.0 with Meeko 0.7.1 and RDKit 2025.9.6 |
| Evidence status | Single documented workflow demonstration, not a benchmark or binding validation |
The selected size32_1 result is visible in the authentic
workflow figure above with the receptor, search box, top pose,
two-dimensional structure, ligand list, terminal states, ranked scores,
and Vina settings in one workspace. The source inputs, prepared files,
pose outputs, scripts, checksums, and machine-readable execution report
are retained with the article's visual-evidence record.
The scores should only be compared within the limits of this protocol. We did not perform target-specific enrichment, alternate-state enumeration, molecular dynamics, free-energy calculations, synthesis, or an activity assay for this demonstration. The useful claim is operational: the workflow accepted six inputs, preserved two failures, completed six docking runs, and kept the evidence connected.
What to evaluate in molecular docking software for generated libraries
Buyer checklist: capability, evidence, and practical outcome
| Question to ask | Evidence to request | Why it matters |
|---|---|---|
| Which ligand formats and batch sizes are genuinely supported? | A real import and preparation run with your representative files. | Generation output must survive the actual handoff, not only a marketing format list. |
| Are preparation failures visible and exportable? | Per-ligand terminal states and diagnostic messages. | Silent loss changes the denominator and weakens traceability. |
| Can source and prepared structures remain distinct? | Linked input, prepared, and output records. | Reviewers need to know what changed before docking. |
| Are receptor and box decisions recorded? | Exact structure identity, preparation record, coordinates, dimensions, and visual context. | The sampled site is part of the scientific question. |
| Can the protocol be reproduced? | Engine version, settings, seeds, statuses, and downloadable outputs. | A detached score cannot explain how it was produced. |
| Can poses be reviewed without moving among several tools? | Real receptor, ligand, pose, grid, and score review in the working interface. | Integrated context reduces handoffs and avoidable identity errors. |
| Can a completed job be found again? | Persistent history, search, reopening, and structured exports. | Repeated screening becomes a cumulative research record instead of a folder reset. |
Where MolNexus fits in this workflow
MolNexus product profile as of July 25, 2026
| Product | MolNexus 0.1.0, local Windows desktop molecular docking software |
|---|---|
| Supported system | Windows 10 or Windows 11, 64-bit |
| Receptor inputs | PDB, mmCIF, or direct RCSB PDB retrieval |
| Ligand inputs | SDF, MOL, MOL2, and PDBQT; one ligand or a batch |
| Workflow | Receptor and ligand preparation, interaction-box definition, docking, pose review, local history, and export |
| Docking engine | AutoDock Vina 1.2.7 |
| Visualization | Integrated Mol* receptor, ligand, search-box, and pose review |
| History and exports | Local SQLite history with CSV, PDB, ZIP, PDF, and Excel outputs available in the current interface |
| Generation scope | MolNexus imports generated molecules; it does not run a generative model |
| License | US$499 one-time purchase; one Windows PC at a time; perpetual use of the purchased version; 12 months of updates |
| Availability | Coming soon; purchase and download are not open yet |
Good fit and requirements beyond the current product
| MolNexus is a practical fit when... | A different or additional system is required when... |
|---|---|
| You already have generated small-molecule files in a supported format. | You need software to train or run the generative model itself. |
| You want a local graphical path from preparation through Vina docking and pose review. | You require macOS, Linux, browser delivery, or simultaneous multi-user access. |
| You are evaluating modest repeated batches and need per-ligand terminal states. | You are screening ultralarge libraries on an HPC or distributed-cloud pipeline. |
| You want receptor, ligand, box, scores, poses, history, and exports in one Windows application. | You require flexible-receptor co-generation, molecular dynamics, ABFE, or another engine outside the confirmed scope. |
| You value local data storage and reopening completed jobs. | Your next decision requires synthesis, biochemical assays, cellular evidence, or another experimental endpoint. |
From generated structures to a decision-ready shortlist
Generative chemistry and molecular docking are complementary layers. A generator expands what can be proposed. A docking workflow asks how accepted molecular representations behave within a declared structural model and search protocol. Machine-learning-guided docking studies have shown how computational prioritization can traverse very large chemical spaces, while later experimental testing remains part of the evidence chain [5]. Other generative workflows explicitly use chemistry filters, docking, more intensive simulation, and bioassays as separate gates [6].
For a research team choosing software, the practical question is therefore concrete: Can this system take our actual generated files, show which ones fail, preserve the protocol, let us inspect poses, and return an auditable shortlist? The answer should be demonstrated with real inputs and outputs.
Frequently asked questions
Build the evaluation chain before the candidate list grows
The best time to define provenance, preparation rules, terminal states, docking parameters, and review criteria is before a generated library becomes difficult to trace. Start with a small representative subset. Measure how many structures enter, prepare, fail, dock, and reach review. Then decide whether the workflow is sufficiently transparent and efficient to scale.
If you are comparing broader software requirements first, use the molecular docking software buyer's guide. If the need is already clear, inspect the real MolNexus interface and documented workflow against your own ligand formats and operating environment.
References
- Wang J, Zhang DY, Budakoti S, Dokholyan NV. A diffusion-based framework for designing molecules in flexible protein pockets Science Advances 12(15):eaeb7045 (2026) DOI: 10.1126/sciadv.aeb7045
- Wang J. Dataset for A Diffusion-Based Framework for Designing Molecules in Flexible Protein Pockets Zenodo, version v2 (2025) DOI: 10.5281/zenodo.17702010 CC BY 4.0 dataset containing the YuelDesign CDK2 molecules used for the documented workflow.
- Kang YN, Stuckey JA. Crystal structure of apo CDK2 RCSB Protein Data Bank, PDB 4EK3 (2013) DOI: 10.2210/pdb4EK3/pdb X-ray diffraction structure at 1.34 Å resolution.
- Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings Journal of Chemical Information and Modeling 61(8):3891-3898 (2021) DOI: 10.1021/acs.jcim.1c00203
- Luttens A, Cabeza de Vaca I, Sparring L, Brea J, Martínez AL, Kahlous NA, et al. Rapid traversal of vast chemical space using machine learning-guided docking screens Nature Computational Science 5:301-312 (2025) DOI: 10.1038/s43588-025-00777-x
- Filella-Merce I, Molina A, Díaz L, et al. Optimizing drug design by merging generative AI with a physics-based active learning framework Communications Chemistry 8:238 (2025) DOI: 10.1038/s42004-025-01635-7
- Center for Computational Structural Biology. Docking multiple ligands with AutoDock Vina AutoDock Vina documentation Official batch-docking documentation for preparing and executing multiple ligands with one configuration.