Pairwise DNA Alignment in Seqqio: Compare Constructs, Gaps, and Sequence Coverage
Compare a reference DNA sequence with substitutions, an insertion, and a shorter fragment. Reproduce Seqqio's alignment scores and learn why identity, gaps, and source coverage must be read together.
This tutorial compares one 16-base synthetic reference with an unchanged copy, a one-base substitution, a one-base insertion, and an eight-base fragment. The sequences are short so every result can be inspected directly. The same decisions matter when you compare longer local constructs: what span was assessed, how gaps were scored, and which original sequence each result belongs to.
Set up an end-to-end DNA comparison
Open Pairwise DNA, select Single pair, and put the reference in input A. Put the unchanged copy in input B for the first run. Select Global · end-to-end, set Gap opening to 10 and Gap extension to 0.5, confirm the displayed matrix is NUC.4.4, then select Align pair.
>reference
ACGTTGCACTGATCGT
>unchanged_copy
ACGTTGCACTGATCGT
The two FASTA entries above describe the inputs; for Single pair, paste or load one record into each side. Both sequences retain their supplied orientation. Seqqio does not automatically try a reverse complement. If a query was supplied from the opposite strand, resolve its orientation explicitly before judging differences.
Gap costs in the form are positive penalties. A gap of length k costs opening + (k - 1) × extension, so a one-base gap costs 10 and a four-base gap costs 11.5. Global mode also penalizes terminal gaps. The official Biopython pairwise-alignment tutorial explains how global and local objectives and affine gap parameters affect an alignment. Keep the settings fixed when comparing this tutorial's scores.
Compare a substitution and an insertion
Keep A unchanged and replace B with each variant below. The substitution changes reference base 12 from A to C. The insertion adds an A to the reference string. These are original software controls, not experimentally characterized constructs.
>one_substitution
ACGTTGCACTGCTCGT
>one_insertion
ACGTTGCAACTGATCGT
| B input | Score | Alignment columns | Identical columns | Gap columns | Identity |
|---|---|---|---|---|---|
| Unchanged copy | 80 | 16 | 16 | 0 | 100% |
| One substitution | 71 | 16 | 15 | 0 | 93.75% |
| One insertion | 70 | 17 | 16 | 1 | 94.12% |
Both input coverages are 100% for these three comparisons. Seqqio's identity is identical columns divided by all alignment columns, including gap columns. The insertion therefore gives 16/17, rounded to 94.12%, rather than 100%. Its alignment was:
A ACGTTGC-ACTGATCGT
B ACGTTGCAACTGATCGT
The gap preserves all bases from both original sequences. Nearby repeated A bases can allow equivalent gap placements, so do not turn one displayed optimum into proof of a uniquely resolved insertion coordinate. The score evaluates a path under the chosen model; it is not a probability of relatedness, an E-value, or experimental evidence.
See why 100% identity can cover only half the reference
Replace B with TGCACTGA, the eight-base segment at reference bases 5–12. Run Global first, then switch only Alignment mode to Local · best matching region and align again.
| Metric | Global | Local |
|---|---|---|
| Score | 17 | 40 |
| Alignment columns | 16 | 8 |
| Identical / gap columns | 8 / 8 | 8 / 0 |
| Identity | 50% | 100% |
| Reference A coverage | 100% | 50% |
| Fragment B coverage | 100% | 100% |
| Reference span, one-based inclusive | 1–16 | 5–12 |
Global includes every original base, padding the shorter B sequence with four terminal gaps on each side. Eight matches contribute 40, and the two four-base gap runs cost 11.5 each, giving 17. Local selects the eight matching bases without those flanks, giving 40. The Local result answers where the fragment fits; it says nothing about the unaligned half of A.
Global
A ACGTTGCACTGATCGT
B ----TGCACTGA----
Local, original A bases 5–12
A TGCACTGA
B TGCACTGA
Coverage measures the original source span included in the selected alignment divided by that source's full length. It is not the fraction of alignment columns containing a match. Read the two source coverages separately when input lengths differ. Comparing full candidate constructs usually calls for an end-to-end question; locating a known partial fragment calls for a regional question.
Treat ambiguous bases as unresolved information
Pairwise DNA accepts supported IUPAC DNA symbols, but a literal symbol match is not necessarily a confirmed base match. In our control, Global alignment of NNNN with NNNN had 100% literal identity and score −4 under NUC.4.4. Those four N symbols do not tell you which concrete bases are present.
For the same N-only inputs, Local returned score 0, an empty alignment, zero coverage, and no defined identity percentage. An empty best region is a valid outcome under this profile. Preserve that distinction instead of reporting “0% identity,” substituting concrete bases, or treating the result as a calculation failure.
Move from one pair to an auditable batch
Once the single-pair controls make sense, choose Batch. Reference A vs every query in B compares one selected A record independently with each B record. A1 vs B1, A2 vs B2, … pairs two equally sized FASTA inputs by original record order. Matching titles do not define paired-batch partners. Confirm order and identifiers before running.
Save the selected pair's aligned FASTA and the dataset report. Keep the original inputs, matrix, mode, gap costs, method profile Seqqio-PairwiseDna-Affine-v1, source coordinates, and both coverage values with your review. Batch pairing organizes separate comparisons; it does not create a multiple sequence alignment.
Use the existing Global Alignment guide and Local Protein Alignment guide for the conceptual background. If your question is instead “where does this full DNA motif occur with a bounded number of edits?”, the Seqqio Fuzzy DNA Search tutorial addresses that different result set.
Choose Seqqio for a declared pairwise question
Pairwise DNA fits local construct review, variant demonstrations, and comparisons between a reference and supplied fragments when an explicit scoring model is appropriate. Local mode returns one optimal region, not all possible matching sites. Use a dedicated reference-search or read-mapping workflow when you need indexed genome search, mapping quality, database significance, or large collections of reads. Per-tool work limits still apply.
Pairwise DNA is included in Seqqio's 39-application Windows 64-bit toolkit. The current offer is a US$99 one-time purchase with no activation key. Review the platform requirements and final checkout total on the product page.
References
- Biopython contributors. Pairwise sequence alignment Biopython 1.88 Tutorial Official global/local and affine-gap context. Independent example execution used installed Biopython 1.87.