Computational Immunology For Classic Antigenicity

How to Predict Protein Antigenicity In Silico: A Reproducible Workflow

Learn what a Kolaskar–Tongaonkar-derived propensity profile measures, how Classic Antigenicity calculates it, and how to prioritize candidate linear regions with reproducible evidence.

Folded protein with highlighted surface regions beside an unlabeled antigenicity propensity profile
Conceptual sequence-based antigenicity screening with candidate regions highlighted for downstream comparison.

A protein may contain hundreds or thousands of residues, while an experimental campaign can test only a fraction of the possible peptides or surface regions. Computational screening helps reduce that search space. The useful question is therefore not “did the software prove that this is an epitope?” but “which regions deserve the next unit of experimental effort, and why?”

This guide uses the Kolaskar–Tongaonkar-derived profile implemented in BioChemIntelli’s Classic Antigenicity application as a transparent first-pass method. The same reporting principles apply to newer sequence- and structure-based predictors.

What an antigenicity profile actually measures

The 1990 Kolaskar–Tongaonkar paper derived a 20-residue antigenic propensity scale from 156 experimentally described segmental determinants containing 2,066 residues. Surface occurrence was estimated from hydrophilicity, accessibility, and flexibility, and overlapping heptapeptide means were assigned to their central residues [1].

Classic Antigenicity uses that published propensity scale, but it is not a line-for-line reproduction of the paper’s algorithm. The application calculates a centered moving average for a user-selected window, shortens the window at sequence termini, applies a configurable threshold that defaults to 1.0, and reports a candidate region after at least seven consecutive profile positions meet that threshold. The paper used complete seven-residue windows, selected either 1.0 or the whole-protein mean as its cutoff, and required a minimum run of six residues [1].

The resulting output is a sequence-derived propensity profile, not a model of a particular antibody or of the complete folded antigen. A large benchmark of amino-acid propensity scales found that even the best scale-and-parameter combinations were only marginally better than random at locating mapped B-cell epitopes [2]. That evidence supports using this profile for transparent triage rather than as a stand-alone classifier.

A reproducible five-step screening workflow

1. Define the biological question

State what you are trying to prioritize before running a predictor. Examples include selecting peptides for an antibody-binding screen, comparing variants, finding exposed regions for reagent development, or generating hypotheses for an immunodiagnostic study. The intended assay determines which evidence matters and prevents a score from being stretched beyond its purpose.

2. Prepare a traceable protein sequence

Use a defined sequence version, not an unlabeled string copied between documents. Record:

  • the accession, database release, isoform, construct, or internal version;
  • whether signal peptides, propeptides, tags, or transmembrane regions were retained;
  • every engineered mutation or strain-specific substitution; and
  • a checksum so the exact input can be reconstructed later.

Plain sequences and single-record FASTA inputs are suitable for Classic Antigenicity. Ambiguous or non-standard residues should be resolved rather than silently replaced.

3. Generate and record a baseline profile

Run the sequence with a documented odd window size and threshold. Classic Antigenicity defaults to a seven-residue window and a score threshold of 1.0. It returns the full residue profile, candidate regions containing at least seven consecutive positions at or above the threshold, and a table of all fixed-length peptides. Odd windows such as 5, 7, or 9 keep a single residue at the center of the local average.

The peptide-table score is the mean of the already smoothed residue profile across each peptide; it is therefore a second aggregation, not the raw mean of the 20 published amino-acid propensity values. Preserve the complete result rather than copying only the highest-scoring segment.

Minimal run record
sequence_id: <accession-or-internal-version>
sequence_sha256: <checksum>
method: Kolaskar-Tongaonkar-derived propensity profile
implementation: BioChemIntelli Classic Antigenicity
window_size: 7
threshold: 1.0
region_min_length: 7
run_date: 2026-07-24

4. Test robustness instead of chasing one peak

Repeat the analysis with nearby odd window sizes—such as 5, 7, and 9—and thresholds close to the observed profile distribution. The published amino-acid scale spans approximately 0.776 to 1.412, and local averaging narrows the attainable range further, so thresholds far outside the observed scores contain no useful discriminatory information.

A region that remains prominent under modest parameter changes is easier to justify than a narrow peak that disappears immediately. This is a sensitivity analysis, not independent validation: every run still uses the same scale and related inputs. Define the comparison in advance and report unstable or negative results alongside prominent regions.

5. Add independent structural and biological evidence

Evidence layers for candidate prioritization

Evidence layer Question answered Main limitation Practical action
Sequence propensity Which local regions score highly under the selected implementation? Does not represent a specific antibody or full 3D context. Rank candidates and document the exact settings.
Structure and accessibility Is the region exposed in a relevant conformation? A model or isolated structure may not capture biological states. Inspect experimental structures or quality-controlled models.
Conservation and specificity Is the region conserved in targets and distinct from off-targets? Depends on representative sequence sampling. Use curated alignments and record the sequence set.
Known epitope evidence Has a matching or overlapping region been tested? Assay, host, and outcome context may differ. Search IEDB, then read and cite the underlying experiment [6].
Experimental validation Does the candidate perform in the intended assay? Results remain assay- and context-specific. Predefine controls, acceptance criteria, and replication.

How to interpret Classic Antigenicity output

  • Profile curve: use it to compare regional trends, not to assign residue-level certainty.
  • Threshold: it is an analysis setting, not a universal biological boundary.
  • Predicted regions: the application requires at least seven consecutive profile positions at or above the threshold.
  • All-peptides table: its score averages the smoothed profile again across each fixed-length peptide.
  • Sequence termini: terminal profile values use fewer residues than central values because the centered window is truncated at each edge.

Use rank, sensitivity to settings, and genuinely independent evidence together. Avoid converting the output into an unsupported binary label such as “antigenic protein” or “confirmed epitope.”

Common interpretation mistakes

  1. Calling a prediction a confirmed epitope. Confirmation requires an appropriate binding or immune-recognition experiment.
  2. Conflating B-cell and T-cell evidence. Antibody recognition, antigen processing, MHC binding, and T-cell recognition are related but distinct questions.
  3. Ignoring conformation. Many antibody epitopes are discontinuous in primary sequence and depend on the folded protein surface [4,5].
  4. Reporting only the best hit. This hides alternative candidates and parameter sensitivity.
  5. Using an unversioned input. Coordinates are meaningless if the underlying construct cannot be recovered.
  6. Selecting a cutoff after seeing the answer. Post-hoc tuning inflates apparent confidence.

A compact reporting checklist

  • Biological question and intended downstream assay
  • Sequence identifier, version, construct boundaries, and checksum
  • Software implementation, method lineage, date, odd window size, threshold, and preprocessing
  • Complete candidate list with coordinates, region length, and scores
  • Parameter-sensitivity results
  • Structural accessibility and conservation or specificity analysis
  • Relevant experimental records and primary citations
  • Decision rule, unresolved questions, and proposed validation experiment

References

  1. Kolaskar AS, Tongaonkar PC. A semi-empirical method for prediction of antigenic determinants on protein antigens FEBS Letters 276(1–2):172–174 (1990) DOI: 10.1016/0014-5793(90)80535-Q
  2. Blythe MJ, Flower DR. Benchmarking B cell epitope prediction: underperformance of existing methods Protein Science 14(1):246–248 (2005) DOI: 10.1110/ps.041059505
  3. Larsen JEP, Lund O, Nielsen M. Improved method for predicting linear B-cell epitopes Immunome Research 2:2 (2006) DOI: 10.1186/1745-7580-2-2
  4. Jespersen MC, Peters B, Nielsen M, Marcatili P. BepiPred-2.0: improving sequence-based B-cell epitope prediction using conformational epitopes Nucleic Acids Research 45(W1):W24–W29 (2017) DOI: 10.1093/nar/gkx346
  5. Kringelum JV, Nielsen M, Padkjær SB, Lund O. Structural analysis of B-cell epitopes in antibody:protein complexes Molecular Immunology 53(1–2):24–34 (2013) DOI: 10.1016/j.molimm.2012.06.001
  6. Vita R, Blazeska N, Marrama D, et al. The Immune Epitope Database (IEDB): 2024 update Nucleic Acids Research 53(D1):D436–D443 (2025) DOI: 10.1093/nar/gkae1092
  7. Ponomarenko J, Bui HH, Li W, Fusseder N, Bourne PE, Sette A, Peters B. ElliPro: a new structure-based tool for the prediction of antibody epitopes BMC Bioinformatics 9:514 (2008) DOI: 10.1186/1471-2105-9-514