Evaluate a Virtual Screen: EF1% and Precision–Recall from a Ranking Table

A favorable docking score does not tell you whether a ranked list retrieves known actives early. Follow a complete numerical example and distinguish early enrichment from whole-ranking metrics.

This guide is for a researcher deciding whether a ranking method deserves further evaluation before a larger screen. The example is deliberately synthetic: 1,000 labeled entries, 20 positives and a declared ranking. It tests the arithmetic and reporting conventions, not MolNexus performance or the ability of a docking score to identify experimentally active compounds.

Keep one consistent row per evaluated molecule

Begin with a stable identifier, the reference label, a score and a documented ranking rule. Decide how multiple protonation states, tautomers, conformers and docking poses are aggregated before calculating molecule-level metrics. Counting many poses of the same active as separate successful molecules can distort both the denominator and the apparent retrieval.

State whether a larger or smaller score ranks better. Vina scores are conventionally ordered from more negative to less negative for this use, while many metric functions expect larger values to indicate the positive class. When supplying a Vina score to such a function, use a consistent transformation such as its negative, after defining the molecule-level aggregation. Do not mix score direction between runs.

First ten ranks in the synthetic example
RanksPositive labelsInterpretation
1, 2, 5, 101Four positives in the top ten
3, 4, 6, 7, 8, 90Six negatives in the top ten

Calculate the early-retrieval quantities explicitly

Total entries N = 1000
Total positives A = 20
Top 1% count k = 10
Positives among the top k, a = 4

Precision@10 = a / k = 4 / 10 = 0.40
Recall@10 = a / A = 4 / 20 = 0.20
Evaluation-set prevalence = A / N = 20 / 1000 = 0.02
EF1% = (a / k) / (A / N) = 0.40 / 0.02 = 20

The top ten entries have a 40% positive fraction, compared with 2% in the complete constructed set. That is 20-fold enrichment at this cutoff. The same result also retrieves only four of the twenty positives: recall is 20%. A strong-looking enrichment value therefore does not mean that most positives have been recovered.

The cutoff must be reproducible. For a set whose size is not divisible by 100, declare whether you round the top-1% count up or down. Prespecify how ties at the cutoff are treated; an arbitrary identifier-based tie break can change early metrics when only a few entries are selected. This example uses 1,000 distinct ranking scores, so neither issue affects its arithmetic.

Reproduce the complete ranking and both PR summaries

The following code constructs every label from the declared positive ranks. All other ranks are negative. Scores are negative rank numbers, so larger values rank earlier; they are not docking energies or predicted probabilities.

import numpy as np
from sklearn.metrics import average_precision_score, precision_recall_curve, auc

positive_ranks = [1, 2, 5, 10, 20, 40, 80, 120, 160, 220,
                  300, 380, 460, 540, 620, 700, 780, 860, 940, 1000]
ranks = np.arange(1, 1001)
labels = np.isin(ranks, positive_ranks).astype(int)
scores = -ranks.astype(float)

precision_at_10 = labels[:10].mean()
recall_at_10 = labels[:10].sum() / labels.sum()
ef_1_percent = precision_at_10 / labels.mean()
ap = average_precision_score(labels, scores)
precision, recall, _ = precision_recall_curve(labels, scores)
pr_auc_trapezoid = auc(recall, precision)
print(precision_at_10, recall_at_10, ef_1_percent)
print(ap, pr_auc_trapezoid)
Executed results with NumPy and scikit-learn 1.7.2
QuantityValue
Precision@100.40
Recall@100.20
EF1%20
Average precision, non-interpolated0.195415
Trapezoidal area under the PR curve0.188547

scikit-learn’s average_precision_score documentation defines average precision as recall-weighted precision. It is not the same calculation as trapezoidal area under a precision–recall curve. Both values above come from the same ranking. Label the metric and implementation explicitly instead of using “PR-AUC” for several different conventions.

Interpret metrics in the population you actually evaluated

Saito and Rehmsmeier’s precision–recall study explains why precision–recall analysis is informative for imbalanced classification. For a screen, precision helps describe the positive fraction in the selected set, but it depends on the evaluation labels and their prevalence. A benchmark enriched with known actives does not directly predict the hit rate in a new experimental library.

Known inactives, presumed negatives and computational decoys are not interchangeable evidence. Define the assay, activity threshold and uncertainty behind each label. Check for duplicate structures and close analog series shared with protocol-development data. Otherwise a favorable ranking may partly reflect a construction artifact or an easier chemical domain than the intended screen.

Keep preparation and docking failures visible. Declare whether failed molecules remain at the bottom, are analyzed separately or are excluded under a prespecified rule, and report the count. Removing failures after seeing their labels changes the evaluation population. For uncertainty estimates, also consider the dependence among chemical series rather than treating every close analog as independent evidence.

Connect the calculation to a screening decision

Choose the early cutoff from a real review or assay budget, compare methods on the same defined set, and reserve an independent assessment from protocol tuning. Report the full ranking, label provenance, aggregation rules, failure counts, early metrics and whole-ranking convention. There is no universal EF1% value that validates a screen for every target.

MolNexus can organize ligand preparation, docking execution, pose review and exported results. The label assembly and enrichment calculations in this article are external analysis steps, not built-in MolNexus enrichment features. Use the validation workflow guide to connect ranking evaluation with preparation and pose-recovery controls.

References

  1. scikit-learn developers. average_precision_score scikit-learn documentation Average precision is recall-weighted precision, distinct from trapezoidal interpolation.
  2. Takaya Saito; Marc Rehmsmeier. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets PLOS ONE (2015) Precision-recall interpretation for imbalanced evaluation sets; not a universal passing threshold.