In Silico Drug Discovery Today: How Structure Prediction, Virtual Screening, AI, and Experiments Work Together

Modern in silico drug discovery is not one algorithm. Learn how target evidence, structure prediction, molecular docking, virtual screening, AI, physics-based refinement, and experiments work as one traceable decision system.

Researchers now have access to accurate protein-structure predictions, ultra-large purchasable libraries, molecular-docking engines, generative models, learned property predictors, and increasingly mature free-energy methods. That abundance creates a new problem: a workflow can look technically advanced while its outputs answer different questions, rest on incompatible assumptions, or lose the provenance needed to reproduce a decision.

This guide maps the present-day workflow from biological question to experimental feedback. It is intended for individual computational researchers choosing tools and for laboratories evaluating whether a workflow can be reviewed, repeated, and handed from one specialist to another.

In silico drug discovery is a decision system, not one model

The modern workflow at a glance

Layer Primary question Useful output What that output does not establish
Target and disease biology Which intervention could change a relevant phenotype? Target hypothesis, mechanism, assay plan That modulating the target will be effective or safe
Structure selection or prediction Which molecular state represents the question? Experimental structure, predicted model, confidence and state annotations That a binding site or ligand pose is correct for the intended experiment
Chemical identity and preparation Which molecular forms will be evaluated? Traceable 3D states with explicit preparation decisions That those states dominate under assay conditions
Docking and virtual screening Which poses and compounds deserve closer inspection? Candidate poses, scores, rankings and failure records Binding, activity, selectivity or developability
Machine learning Which patterns learned from data may help prioritize candidates? Predictions with an applicability domain and uncertainty Reliable behaviour outside the model's evidence domain
Physics-based refinement How might related candidates compare under a defined physical model? Relative or absolute energetic estimates with diagnostics A universal, assumption-free measure of affinity
Experiments Does the hypothesis hold in the selected assay system? Measured response, controls, failure modes and new training evidence Automatic translation across assays, models, organisms or patients

1. Begin with biology and a decision question

The first computational input should not be a protein file. It should be a decision question. Are you looking for ligands that occupy a known orthosteric site, stabilize a particular conformation, disrupt an interaction, explain a resistance mutation, or create a chemically diverse set for an assay?

That question determines the relevant target state, controls, library composition, computational methods, and experimental endpoint. Without it, a workflow can optimize an attractive score against a biologically irrelevant state. Record the target rationale, intended mechanism, assay context, known liabilities, and the criterion that would justify the next experiment before screening begins.

2. Select or predict the relevant molecular structure

AlphaFold 2 changed access to protein-structure hypotheses by achieving high accuracy in the CASP14 assessment and reporting confidence estimates that help users identify uncertain regions [1]. AlphaFold 3 broadened structure prediction to complexes that can include proteins, nucleic acids, small molecules, ions, and modified residues [2]. These advances reduce a historic barrier: many projects no longer have to begin with no structural model at all.

Structure availability is not the same as structure suitability. A docking receptor still needs a relevant biological state, credible binding-site geometry, appropriate cofactors or waters, repaired chemistry, and an explicit preparation policy. Confidence can vary locally even when the global fold looks convincing. A predicted complex is also not interchangeable with an AutoDock Vina-style prospective screen of a user-defined ligand library.

Use the AlphaFold and molecular-docking guide when deciding how a predicted receptor can enter a small-molecule workflow without overstating what structure prediction has established.

3. Define chemical identity before screening

A database row or SMILES string is not yet a docking-ready scientific object. Salt handling, protonation, tautomerism, stereochemistry, valence, explicit hydrogens, three-dimensional geometry, atom typing, and charge assignment can change the object evaluated by the next method.

Keep the source identity separate from every prepared state and retain failures rather than allowing unsupported molecules to disappear from the denominator. The protein and ligand preparation guide explains the scientific choices; the ligand preflight guide shows how to detect structural problems before they become ambiguous docking failures.

4. Use docking and virtual screening to prioritize hypotheses

Structure-based virtual screening asks whether a search method can generate plausible poses and rank compounds usefully enough to reduce a much larger library to a smaller set for review or testing. Prospective studies have shown that docking can operate at enormous scale and recover previously unrecognized chemotypes, as demonstrated in screens of more than 170 million make-on-demand compounds [3] and in the open-source VirtualFlow platform for ultra-large screens [4].

Scale is not proof of accuracy. A defensible screen preserves the receptor, ligand states, search box, scoring function, exhaustiveness, random seed policy, software version, failures, ranked poses, and review decisions. It validates the protocol with target-relevant evidence before exposing the production library to it.

What docking can and cannot answer

Docking can support Docking alone cannot establish
Generating candidate binding poses That the pose occurs in the experimental system
Prioritizing compounds under one controlled protocol That rank order equals measured potency
Exploring interactions and steric compatibility Binding kinetics, cellular activity, selectivity, or safety
Reducing a library to a reviewable or testable set That excluded compounds are inactive

5. Use AI where the evidence domain matches the question

Machine learning can prioritize compounds, estimate properties, generate molecular proposals, predict poses, and identify patterns across data that are difficult to encode manually. Its role is strongest when the training data, representation, endpoint, and deployment domain match the decision being made.

There are persuasive examples of computational prediction connected to experiments. Stokes and colleagues trained a deep learning model on antibacterial measurements, applied it to other chemical collections, and experimentally validated prioritized compounds [8]. More recently, Scalia and colleagues combined a primary screen of roughly two million compounds with deep learning and prospective testing after screening more than 1.4 billion synthesizable compounds in silico [7]. These are integrated discovery studies, not evidence that a generic model works for every target, assay, or chemical space.

6. Use physics-based refinement for narrower decisions

Molecular dynamics and free-energy calculations address questions that a docking score was not designed to answer. They can examine conformational stability, interactions over time, or energetic differences among a narrower set of compounds. In an industry-scale assessment, Schindler and colleagues evaluated relative binding free-energy calculations across active drug discovery projects and documented both prospective usefulness and practical constraints [5].

These methods are not automatic confirmation layers. Their reliability depends on the structural model, chemical mapping, force field, sampling, protocol, diagnostics, and the type of chemical change. They are usually most defensible when the question has narrowed enough to justify their greater setup and compute cost.

7. Make developability a multi-objective decision

A compound can have an attractive predicted pose and still be a poor experimental candidate. Solubility, permeability, chemical stability, aggregation risk, reactivity, metabolism, selectivity, synthesis, and assay compatibility can change the value of a shortlist.

Use predicted properties to expose trade-offs and prioritize measurements, not to collapse every objective into an unexplained composite score. Preserve the individual values, prediction methods, confidence or domain indicators, decision thresholds, and the reason a candidate advanced. The best candidate is the one that fits the current decision, not necessarily the molecule with the most extreme value in any single column.

8. Treat experiments as the feedback loop

Experiments do more than approve the final shortlist. They expose assay interference, target dependence, chemical-state errors, inactive regions of chemical space, and model blind spots. Those results should return to the computational workflow as new evidence rather than being reduced to a success/failure headline.

Grisoni and colleagues demonstrated this loop by coupling generative molecular design with automated synthesis and in vitro testing [9]. The antibacterial studies above likewise connected computational prioritization to prospective measurements [7,8]. The transferable principle is not a particular model: design the computational and experimental records so that each tested compound can improve the next decision cycle.

A defensible handoff between workflow layers

Handoff Minimum record to preserve
Biology to structure Target hypothesis, biological state, construct, species, cofactors, mutations and assay context
Structure to docking Source or model version, confidence, preparation, binding-site rationale, retained waters and receptor artifact
Library to prepared ligands Source ID, original record, enumerated state, preparation policy, prepared artifact and failure reason
Docking to review Engine version, search box, settings, seed policy, scores, poses, logs, failures and review annotations
Prediction to experiment Model version, training boundary, uncertainty, selection rule, synthesized or purchased identity and assay protocol
Experiment to next cycle Raw and processed measurements, controls, exclusions, decision outcome and model-ready labels

How the layers work together in practice

  1. Frame the decision. Define the biological question, target state, assay, controls, and advancement rule.
  2. Choose structural evidence. Select an experimental structure or predicted model and document local confidence and state relevance.
  3. Curate the chemical set. Preserve source identities, enumerate intended states, and retain preparation failures.
  4. Validate before screening. Test whether the docking protocol is fit for the intended prioritization task.
  5. Generate and review hypotheses. Dock under a frozen protocol, inspect poses, and keep score, geometry, and provenance together.
  6. Add orthogonal evidence selectively. Use learned property models, similarity, physics-based refinement, or other methods when each answers a defined unresolved question.
  7. Select a diverse, testable set. Balance predicted evidence, chemical diversity, liabilities, synthesis or availability, and assay capacity.
  8. Measure and learn. Return experimental outcomes, including failures and controls, to the next cycle.

This sequence is a framework, not a mandatory software stack. Some projects begin with phenotypic data, fragment hits, known ligands, or an experimental structure. The invariant requirement is that each method has an explicit role and a traceable handoff.

How to evaluate in silico drug discovery tools

Questions to ask before adopting a tool

Evaluation area Question
Scientific scope Which decision does the tool support, and which claims remain outside its scope?
Input state What chemistry, structure, metadata, and preparation does the result assume?
Method identity Are engine, model, database, parameters, and versions visible and exportable?
Validation What target-relevant controls, prospective evidence, and applicability limits are available?
Failure handling Are rejected inputs and failed calculations preserved with reasons?
Traceability Can a result be reconstructed from source input through review and export?
Interoperability Can structures, tables, logs, and metadata move to the next method or colleague?
Operational fit Do platform, compute, privacy, licensing, maintenance, support, and purchasing terms fit the user or laboratory?

An individual researcher may prioritize rapid setup, local control, transparent files, and a manageable one-time cost. A laboratory or company may additionally need standardized protocols, reviewable histories, deployment rules, purchasing clarity, continuity, and reliable handoffs. The scientific core should remain the same; the operational proof required by each buyer is different.

Where MolNexus fits—and where it does not

MolNexus is focused on one part of this system: reproducible protein-small-molecule docking on a local Windows desktop. Its current product specification connects receptor and ligand preparation review, interaction-box definition, AutoDock Vina execution with Vina or Vinardo scoring, pose inspection, local job history, and exports.

That scope matters. MolNexus should be evaluated as a structured docking workflow, not as target discovery, a generative chemistry system, a free-energy platform, or experimental validation. It can make the docking layer easier to operate and audit; it does not turn a docking result into a biological conclusion. The commercial release is coming soon, and purchase and download are not yet open.

MolNexus in the broader workflow

MolNexus supports The surrounding project must still provide
Prepared receptor and ligand review Target rationale, relevant biological state, and chemical-state policy
Interaction-box and docking-parameter control Target-relevant protocol validation and decision thresholds
Vina or Vinardo execution and ranked poses Independent evidence for binding, activity, selectivity, and developability
Pose review, local history, and exports Downstream modelling, experimental testing, and feedback into the next cycle

Frequently asked questions

The practical conclusion

The present era is defined less by one winning algorithm than by better connections among biology, structures, chemical data, docking, learned models, physical simulations, and experiments. Structure prediction has made credible starting models available to many more projects. Virtual screening can explore chemical spaces at remarkable scale. AI can learn useful patterns and propose candidates. Physics-based methods can refine narrower comparisons. None removes the need to state the question, preserve provenance, test applicability, and measure the result.

A strong in silico workflow makes each handoff explicit. That is what turns a collection of modern tools into a reproducible discovery process.

References

  1. Jumper J, Evans R, Pritzel A, et al.. Highly accurate protein structure prediction with AlphaFold Nature (2021) DOI: 10.1038/s41586-021-03819-2 Original AlphaFold 2 research article supporting the discussion of protein-structure prediction accuracy, confidence, and expanded access to structural hypotheses.
  2. Abramson J, Adler J, Dunger J, et al.. Accurate structure prediction of biomolecular interactions with AlphaFold 3 Nature (2024) DOI: 10.1038/s41586-024-07487-w Original AlphaFold 3 research article supporting the bounded claim that structure prediction now covers several classes of biomolecular complexes.
  3. Lyu J, Wang S, Balius TE, et al.. Ultra-large library docking for discovering new chemotypes Nature (2019) DOI: 10.1038/s41586-019-0917-9 Original prospective study used to support docking at ultra-large scale and discovery of new chemotypes, without generalizing its performance to other targets.
  4. Gorgulla C, Boeszoermenyi A, Wang ZF, et al.. An open-source drug discovery platform enables ultra-large virtual screens Nature (2020) DOI: 10.1038/s41586-020-2117-z Original VirtualFlow study supporting the feasibility and infrastructure of open-source ultra-large virtual screening.
  5. Schindler CEM, Baumann H, Blum A, et al.. Large-Scale Assessment of Binding Free Energy Calculations in Active Drug Discovery Projects ChemRxiv author preprint; peer-reviewed in Journal of Chemical Information and Modeling (2020) DOI: 10.26434/chemrxiv.11364884.v2 Original author preprint for the large-scale assessment. The peer-reviewed article DOI is 10.1021/acs.jcim.0c00900.
  6. Buttenschoen M, Morris GM, Deane CM. PoseBusters: AI-based docking methods fail to generate physically valid poses or generalise to novel sequences Chemical Science (2024) DOI: 10.1039/D3SC04185A Original benchmark supporting physical-validity and generalization checks for the assessed deep-learning docking methods; it is not used as a claim about every AI method.
  7. Scalia G, Seiler M, Steger M, et al.. Deep-learning-based virtual screening of antibacterial compounds Nature Biotechnology (2025) DOI: 10.1038/s41587-025-02814-6 Original prospective study connecting primary-screen data, deep learning, ultra-large virtual screening, and experimental validation in a defined antibacterial setting.
  8. Stokes JM, Yang K, Swanson K, et al.. A Deep Learning Approach to Antibiotic Discovery Cell (2020) DOI: 10.1016/j.cell.2020.01.021 Author manuscript in PubMed Central for the original study connecting learned antibacterial predictions to prospective experimental testing.
  9. Grisoni F, Huisman BJH, Button AL, et al.. Combining generative artificial intelligence and on-chip synthesis for de novo drug design Science Advances (2021) DOI: 10.1126/sciadv.abg3338 Full original article in PubMed Central supporting an integrated example of generative design, automated synthesis, and in vitro testing.