Help

Usage notes, input formats, interpretation of outputs, and troubleshooting across all tools.

Quick start

  1. Open a tool from the menu (e.g., Primer Design, KASP, LAMP, Gibson Assembly).
  2. Provide input:
    • Paste sequences in FASTA format, or
    • Upload a FASTA file, or
    • Retrieve sequence by an accession/variant ID (where supported).
  3. Adjust parameters (primer length, Tm, product size, assay options).
  4. Click the tool's run button (Generate, Design Primers, Design Probes, …). Results appear in the output tabs — copy/paste to Excel when needed.

Input formats

FASTA

Use standard FASTA with a header line starting with > and one or more lines of sequence:

>Seq1
ACGTTGCAACGTTGCAACGTTGCA
>Seq2
TTTACCGGAACTGACTGACT

Name + sequence (PrimersList)

The PrimersList tool also accepts space- or tab-separated name sequence pairs (Excel-friendly):

m13-47  cgccagggttttcccagtcacgac
RP      tttcacacaggaaacagctatgac

Allowed nucleotide codes

Standard and degenerate bases are accepted using IUB/IUPAC codes:

  • N=A/C/G/T, R=A/G, Y=C/T, S=G/C, W=A/T
  • K=G/T, M=A/C, B=C/G/T, D=A/G/T, H=A/C/T, V=A/C/G
  • U=Uracil, I=Inosine
Modified bases (PrimerAnalyser / PrimersList)
LNA nucleotides: E=LNA-dA, F=LNA-dC, J=LNA-dG, L=LNA-dT. Use only where explicitly documented on the tool page.

Targeting markup inside sequences

Several tools allow you to constrain primer placement directly inside the sequence by adding markup characters.

Primer targeting brackets

Square brackets ([, ]) mark regions where primer binding or variants should be evaluated.

...ACCTG [A/T] GGTCA...
Excluded regions

Use /.../ to exclude regions from primer placement (can be repeated multiple times).

...ACCTG /REPEAT/ GGTCA...

Exact interpretation depends on the specific tool. When in doubt, keep the input unmarked and use parameter filters.


Common parameters & output

These settings and output columns behave the same way across the primer/probe design tools. Each tool page lists only the values that differ (defaults and any extra options); the meaning is defined once here.

Design parameters

  • Primer length / Tm range — the design window for the annealing part of each primer. Keep a narrow Tm interval for multiplex compatibility. Tool defaults differ and are listed on each tool.
  • Tm calculation conditions — every design tool computes Tm under the same fixed reaction conditions (55 mM monovalent salt, 1 mM Mg²⁺, 0.2 µM primer — a standard-PCR profile), regardless of the assay's real chemistry. This is deliberate: LAMP, for example, normally runs with much higher Mg²⁺ and primer concentrations, which would raise Tm and make values non-comparable across tools. One fixed condition set keeps Tm numbers comparable across the design tools — treat them as a consistent design-time reference, not a literal prediction of each method's actual reaction conditions. The exceptions are PrimersList and PrimerAnalyser, where salt, Mg²⁺ and oligo concentration are editable on the page; leave them at their defaults (55 mM / 1 mM / 0.2 µM) to keep those Tm values comparable with the design tools.
  • Product / amplicon size — match the downstream method (short for qPCR / RPA, longer for Sanger).
  • 3′-end composition pattern — constrains the last bases of a primer to control specificity: N = any; one or more IUPAC letters; or several equal-length patterns separated by spaces (e.g. sws ssw sww wss www). Worked example in PCR primer design → 3′-end composition variants.
  • Linguistic complexity (LC%) and purine/pyrimidine balance (YR%) — reject low-complexity / repetitive oligos (≥70% LC recommended). Full definition in PCR primer design → Linguistic sequence complexity.

Mask repeats & low-complexity regions

When enabled (on by default in most tools), repeat-rich and simple-sequence tracts are excluded from primer / probe placement — this avoids the most common cause of design failure. The masking uses the same engine as TotalRepeats; you can also pre-mask a sequence there and paste it back in.

Sequence retrieval by ID

Where supported, sequences can be fetched automatically instead of pasted: by rsID from Ensembl (select species and flank size) or by accession from NCBI. If retrieval fails, see Troubleshooting → Sequence retrieval fails.

Standard output columns

Most tools report per-oligo properties in the same columns:

ID · Sequence · Length · Tm · CG% · LC% · YR%

Primer / probe pairs or sets add the relevant Fragment size (bp) and Annealing Tm (°C). Results are tab-separated — copy and paste straight into Excel; see Exporting results. A few tools depart from this list: KASP omits YR%; Universal PCR adds a 3′-conserved and a degenerate-position column; the MLPA and STR / SSR reports head their columns Element … Position, and MLPA has no YR%; and Gator BLI reports only ID · LC% · YR% · Sequence, with no length, Tm or CG%.


PCR primer design

Applies to: PCR / Multiplex / qPCR / RPA

  • Primer length and Tm range: define the design window; keep primer pairs within a narrow Tm interval for multiplex. Tm is computed under fixed standard-PCR conditions for consistency across tools (see Common parameters → Tm calculation conditions).
  • Product size range: match the downstream method (e.g., short for qPCR/RPA; longer for Sanger).
  • Minimal linguistic complexity: fixed at 70% on this tool and not user-settable — primers below that complexity are rejected automatically, which keeps them out of low-complexity / repetitive regions and reduces spurious priming.
  • Non-specific priming control: enable to reduce off-target binding in repetitive genomes.
  • Low complexity priming control: enable to keep primers out of SSR/microsatellite and telomere-like low-complexity stretches.

  • Multiplex PCR: favors primer sets with reduced cross-dimers; use tighter Tm bounds.
  • TaqMan / MGB probe assays: enable one option at a time; probes are designed within the amplicon under assay-specific constraints.
  • RPA: shifts to longer primers and short amplicons suitable for ~37–42 °C isothermal workflows.
  • Inverse PCR: assume circular template; ensure the marked region is consistent with circular amplification logic.
  • C>>T bisulfite conversion: converts non-CpG cytosines for in silico evaluation on bisulfite-treated sequences.
  • Overlapping primers: relaxes constraints on overlap where necessary (use cautiously).

Sequences use the standard IUB/IUPAC nucleic-acid codes. The full list of accepted letters (including U, I and LNA bases) is in Input formats → Allowed nucleotide codes.

The structure of the last nucleotides at the 3′-end of the primer can be specified to control primer specificity:

  • N — any pattern (no constraint)
  • One, two, or more characters using standard or mixed letters
  • Multiple patterns of equal length separated by spaces: sws ssw sww wss www

For example, WSS corresponds to all 3′-end variants: acc acg agc agg tcc tcg tgc tgg.

LC% (Linguistic Complexity)

Measures the "vocabulary richness" of a genetic text by counting nucleotide combinations relative to the theoretical maximum. 100% = highest possible level.

YR% (Purine-Pyrimidine Complexity)

Measures "harmony" of genetic text by counting possible purine-pyrimidine combinations relative to the theoretical maximum. 100% = maximum possible level.

The pre-designed primers/probes list is used for multiplexing with prior designed PCR primer/probe sets. This allows you to incorporate existing validated assays into new multiplex panels while ensuring compatibility and avoiding cross-reactivity.

Design of specific PCR primers for in silico bisulfite conversion for both strands. Only cytosines not followed by guanine (non-CpG context) will be replaced by thymines. CpG methylation sites are preserved.

Optimal primers should hybridize only to the target sequence. Common problems include annealing to repetitive sequences (retrotransposons, transposons, inverted tandem repeats), alternative product amplification, and multiple bands due to off-target binding.

Masks stretches that are poor targets for primer design because of their low sequence complexity — simple/microsatellite tandem repeats (SSRs) and telomere-like repetitive motifs, plus regions similar to them — the same low-complexity blocks detected by Repeat Finder. Such regions are prone to polymerase slippage, unstable annealing, and non-specific priming, so primers are not placed inside them.

This complements Non-specific priming control, which masks dispersed/interspersed repeats found elsewhere in the sequence: low complexity control targets local, in-place repetitiveness at the candidate site itself. Enabled by default.

Minor Groove Binders (MGBs) selectively bind non-covalently to the minor groove of the DNA helix, enabling shorter probe lengths and superior quenching for highly specific TaqMan-type assays.

Results in the PCR Primer Pairs tab are not listed as simple one-to-one forward/reverse pairs. Instead, the output is organised by forward primer: for each forward primer candidate, all compatible reverse primers that pass the cross-dimer, Tm, and product-size filters are listed beneath it. When a probe mode is active (TaqMan or MGB — this tool has no molecular-beacon option), every compatible probe found within the amplicon window of that F+R combination is printed once, on its own row immediately above the reverse-primer rows it belongs to.

This structure lets you:

  • See at a glance how many workable reverse primers exist for a given forward primer — useful when one position has a narrow design window on one strand.
  • Compare amplicon sizes and annealing Tm across all valid F+R combinations and pick the one that best fits your protocol.
  • Choose, per F+R pair, the probe with the most favourable placement, Tm, or GC% without re-running the design.

Reading the output

  • A forward primer row is marked F in the ID column and contains the primer sequence, length, Tm, CG%, LC%, and YR%.
  • Each reverse primer row beneath it is marked R and additionally shows the amplicon size (bp) and the predicted annealing Tm (°C) for that specific F+R pair.
  • If probes are designed, each probe row is printed immediately above the reverse-primer rows it applies to, at the same indentation as the primer rows. Its probe type and position are carried in the ID itself — e.g. 1F_TaqMan_240-262 or 1R_MGB_318-333.
  • When no compatible reverse primer exists for a given forward primer, that forward primer is omitted from the output entirely.

Universal PCR (uPCR) — consensus primers from multiple sequence alignment (MSA)

Applies to: Universal PCR (uPCR)

Universal (consensus) PCR uses a single primer pair to amplify a homologous locus across many divergent templates — related genes, alleles, genotypes, strains or species. Instead of designing against one sequence, the tool aligns two or more DNA sequences (MUSCLE-JS progressive MSA), builds a per-column conservation profile, and places primers so that they bind regions shared by all inputs. Where templates differ, variable positions can be encoded as IUPAC-degenerate bases.

  • Pan-specific detection — one assay that amplifies all members of a family (e.g., a viral genus, a multigene family, paralogues/orthologues).
  • Cross-genotype / cross-strain amplicons — diagnostics and barcoding where targets vary but share conserved anchors.
  • Conserved-region sequencing — generate amplicons spanning a region conserved across all inputs for Sanger/NGS.
  • Requires ≥2 DNA sequences; the more representative the input set, the more robust the resulting universal primers.

  • Paste or upload multi-FASTA DNA (no [brackets] markup is needed — the whole alignment is the target).
  • Auto-correct orientation reverse-complements any read on the opposite strand before alignment, so mixed-orientation inputs still align.
  • The alignment is computed in the browser; gap-open / gap-extend penalties tune how readily indels are introduced.

The Alignment & Conservation tab reports the alignment length, % conserved columns, the list of conserved blocks, and a consensus track (* = 100% conserved, lowercase IUPAC = variable, - = gap column).

Polymerase extension starts at the primer's 3′ terminus, so a mismatch there abolishes amplification. The tool therefore enforces:

  • 3′-end window 100% conserved — the last N bases at the priming end (the “Conserved 3′-window”, default 12 nt) must be identical in every input sequence. This is the minimum guarantee for every primer reported.
  • 5′-end may vary — with “Allow degenerate 5′-end”, upstream positions can sit on variable columns and are written as IUPAC codes (e.g., R=A/G, Y=C/T, N=any). With the option off, the whole primer is restricted to 100% conserved columns (plain A/C/G/T, no degeneracy).
  • Gap columns are never spanned, so each primer has the same length in every template.

This mirrors the standard practice for consensus/universal primers, where the 3′ region is required to have zero variable sites while limited degeneracy is tolerated toward the 5′ end.

  • Length range / Tm range — define the primer search window. Tm is computed on the majority-base (representative) sequence.
  • Product size (bp) — filters pairs by amplicon length; reported as the range of ungapped lengths across all templates.
  • Conserved 3′-window (nt) — number of fully conserved bases required at the 3′ end. Default 12; the field's minimum is 6, and the value never exceeds the minimum primer length.
  • Min. linguistic complexity (%) — rejects low-complexity / repetitive primers (see LC%).
  • Max. pair Tm difference (°C) — keeps forward and reverse primers thermally matched.
  • Allow degenerate 5′-end — toggles IUPAC encoding of variable 5′ positions (see the conservation rule above).
  • Report overlapping primers — when off (default), clusters of overlapping primers are collapsed to the single best candidate by linguistic complexity (LC%), so the output is not flooded with near-duplicate primers whose 3′-anchors differ by only a nucleotide or two. Enable it to list every candidate, including overlapping ones.
  • Max forward sets to report — caps the number of forward-primer blocks shown. Output is grouped: each forward primer is printed once with all its compatible reverse primers (each with product size and Ta) beneath it; blocks are separated by a blank line. The report shows totals vs. displayed.

  • Alignment & Conservation — summary, conserved-block table and consensus/conservation track.
  • Universal Primers — every candidate (Forward/Reverse) with sequence, length, Tm, CG%, LC%/LC_RY%, a 3′-conserved flag and a Degenerate flag.
  • Primer Pairs — compatible F+R sets, each primer with its LC%/LC_RY%, plus product size (bp range) and predicted annealing Ta (°C).

Coordinates are 1-based alignment columns (e.g., F_1-20), so they map directly onto the consensus shown in the conservation track. Use the Copy / Download buttons to export each tab as TSV.

  • No pairs found? Lower the conserved 3′-window, widen the Tm/length limits, or enable a degenerate 5′-end. Very divergent inputs may have no conserved block long enough for a primer.
  • Keep degeneracy modest — highly degenerate primers lower effective concentration and specificity; prefer placing degenerate positions only at the 5′ end.
  • Garbage in, garbage out — misaligned or unrepresentative input sequences yield misleading “conserved” blocks. Inspect the alignment first (the MSA tool helps).
  • This is an in-browser consensus-primer designer; validate candidates by in silico PCR against your real templates before ordering.

Background reading on universal / consensus primer prediction from alignments:


Genotyping (KASP / AS-PCR)

Applies to: KASP primers assay design

The tool computes primers for Kompetitive Allele Specific PCR (KASP) or Allele-Specific Quantitative PCR (ASQ). One Allele-Specific Primer (ASP) is computed for each input allelic variant, plus one common primer (Universal Primer, UP) targeting a conserved region.

Input sequences should be in FASTA format with the variant of interest enclosed in [square brackets]. Supported formats:

  • [First allele/Second allele] — e.g., [A/G]
  • [First/Second/Third/Fourth] — for multi-allelic variants
  • [IUPAC code] — e.g., [R] for A/G, [S] for G/C
  • [Target Nucleotide] — single nucleotide if only one allele needs targeting

Keep flanking regions long enough to allow design of allele-specific primers away from problematic local repeats.

>1
gctctctgtgtctgatccaagaggcgaggccagtttcatttgagcattaa[A/G]tgtcaagttctgcacgctatcatcagggg

>2
tcatattccagtttgggcgagttttaagataggtccgg[S]acagtctttgcggcgccaacgcgtctttctccag

[S] represents G/C using the IUPAC ambiguity code.

Leave one side empty for deletions:

>1
tcatattccagtttgggcgagttttaagataggtccgg[AG/]acagtctttgcggcgccaac

>1
tgggcagcattagtagaagaaagtacaagaccgtgtgtagagg[GATATACTTGAG/CAGTCC]agcagatagcgttggatag

The [square brackets] should surround all SNPs that are part of the haplotype. Nearby SNPs not part of the haplotype should be outside brackets and identified using an IUPAC code.

ASP primers can carry standard KASP FAM and HEX tails (LGC Biosearch Technologies). Paste tails in FASTA-like format in the dedicated tab:

>FAM
GAAGGTGACCAAGTTCATGCT
>HEX
GAAGGTCGGAGTCAACGGATT

Custom tails can be used instead of the standard sequences.

  • Non-specific priming control (default on) — keeps both the allele-specific and common primer off dispersed/interspersed repeats elsewhere in the flanking sequence.
  • Low complexity priming control (default on) — keeps primers off other SSR/microsatellite and telomere-like stretches in the flanks, the same blocks detected by Repeat Finder. Disable either control only if it prevents a primer from being found and the flanks cannot be replaced.

Enter one or more rsIDs (space/comma separated), select a species, and set the desired flank size. The tool retrieves the flanking sequence from Ensembl and populates the input area automatically.

A second panel, Retrieve flanking sequence from NCBI by rsID, does the same job through dbSNP for human variants: enter the rsIDs in the Human rsIDs field, set the flank size and click Retrieve from NCBI. Use it when Ensembl is unavailable or the rsID is not in the selected Ensembl species.


qPCR genotyping (TaqMan / MGB / Molecular beacon / UniQ)

Applies to: SNP/InDel Genotyping — TaqMan / MGB / Molecular beacon Probe Design

The tool designs allele-discriminating probe–primer sets for quantitative PCR (qPCR) genotyping of SNPs and small insertions/deletions. For each variant, a PCR primer pair is computed plus a probe covering the polymorphic position — dual-labelled (TaqMan, MGB or Molecular beacon), or a UniQ pair in which the quencher rides on a separate protector strand shared by all the alleles of that variant. Sequences can be pasted as FASTA, uploaded from a local file, or retrieved directly from Ensembl by rsID. All primers and probes are screened for Tm, hairpins, primer–dimer interactions, repeat masking and linguistic complexity. A Multiplex PCR option designs assay sets for several variants that are mutually dimer-compatible, so they can all be genotyped together in a single reaction — see Primer design parameters and Output tabs and columns below.

  • Paste FASTA in the Input Sequences tab with the polymorphism enclosed in square brackets (e.g., [A/G], [ATCG/-], [R]).
  • Upload a FASTA file — plain FASTA, multi-FASTA or sequences already containing [REF/ALT] brackets.
  • Retrieve from Ensembl by rsID — choose a species from the drop-down (Human, Mouse, Rat, Cow, Pig, Chicken, Arabidopsis), enter one or more rsIDs separated by spaces, set the flank size (≥100 bp) and click Retrieve from Ensembl. The forward-strand context with the allele bracket is fetched automatically.
  • Custom species — type any Ensembl species id (lowercase, underscores, e.g. canis_lupus_familiaris) into the "Other species" field and click Add. Custom species are stored in your browser (localStorage) for later use.
  • Retrieve from NCBI by rsID — a second panel that goes to dbSNP instead of Ensembl, for human variants: enter one or more rsIDs in the Human rsIDs field, set the flank size and click Retrieve from NCBI. Use it when Ensembl is unavailable or the rsID is not in the selected Ensembl species. The retrieved records replace whatever is in the input box.

Provide ≥250 bp of flanking sequence on each side of the variant so the primer search has enough room.

  • Biallelic SNP: flank [A/G] flank or single IUPAC code flank [R] flank
  • Tri- / tetra-allelic SNP: [A/C/G], [A/C/G/T] or [N]
  • Insertion / deletion: use - for the empty side, e.g. [ATCG/-] (4-bp deletion) or [-/T] (single-base insertion)
  • Excluded region — wrap with /.../ to keep primers out; see Targeting markup.
>rs1801133
tggggggaaaattagaggtaaccaaaatgggg...gatgaaatcg[Y]ctcccgcagacaccttctcc...

>rs3918290
gccacatacagtgaaaaccaactcaataaa...aatcacactta[ATCG/-]gttgtctggaaagtcagcc...

All standard IUPAC ambiguity codes are accepted both inside and outside the brackets — see the Input formats section.

Exactly one probe format may be active at a time — checking one option automatically clears the others.

TaqMan (18–28 nt)

Classic 5′-nuclease hydrolysis probe with a 5′ reporter dye and 3′ quencher. Best general-purpose option for short amplicons.

MGB (12–28 nt)

Minor Groove Binder conjugate. Shorter probes with higher Tm and better single-mismatch discrimination — preferred for AT-rich targets and tightly clustered SNPs.

Molecular beacon (18–35 nt)

Hairpin-shaped probe with a stem that quenches the reporter in the absence of target. Extra specificity through structural transition — useful when allele discrimination by hydrolysis probes is borderline.

UniQ (18–30 nt + tail)

Two-strand displacement probe: the SNP sits in the middle of the probe, the reporter on its 5′ end, and the quencher travels on a separate protector strand shared by every allele probe of the locus. No probe is dual-labelled.

How a UniQ set works

The format follows the two-strand toehold-exchange probe of Zhang, Chen & Yin (2012): a complement strand C carrying the signal, held in a duplex by a protector strand P that the target has to compete away.

  • Probe (one per allele). 5′-reporter, then a shared inert tail of 10 nt, then the target-binding region with the variant base in the centre. The 3′ end is blocked (3′-phosphate or C3 spacer) so the probe cannot prime.
  • Protector (one per locus and direction). The reverse complement of the tail plus the 5′-flank the allele probes share — the flank is identical whatever the allele, which is what lets a single strand serve them all. Its 3′-quencher comes to rest against the probe's 5′-reporter, so the reagent is dark before any target appears.
  • Readout. The amplicon binds the free 3′ half of the probe and branch-migrates through the protected flank. The protector is left holding the tail alone, too little to stay on at the annealing temperature, and leaves — the reporter lights up. A centred mismatch stalls that migration where it starts, so the wrong allele stays quenched.
  • The tail. Inert by design: it has no counterpart in the template, and the tool picks it from a fixed set, skipping any that would dimerise with the oligos already in the assay. Its only job is to make the protector duplex stable enough to quench at reaction temperature — a protector covering the short 5′ flank alone would melt off around 30 °C.

Both numbers are printed, and they are the ones to read: the probe Tm is for the probe–target duplex, the protector Tm for the probe–protector duplex. The design keeps the protector at least 4 °C below the probe so the target can win; over a GC-rich flank the protector is trimmed back from the variant to get there, which is also why it is sometimes shorter than the flank it is named for. If the protector still comes out below about 50 °C the report says so — expect background from free probe, and consider the other direction or a different amplicon.

  • Length range (nt) — default 20–22. Keep the window narrow for uniform amplification across loci.
  • Tm range (°C) — default 58–60. Match it to your master-mix recommendation. Probe Tm is internally targeted ~5–10 °C above primer Tm. Computed under the site-wide standard-PCR conditions (see Common parameters → Tm calculation conditions).
  • PCR product size (bp) — default 100–200, already inside the ≤200 bp window qPCR efficiency prefers.
  • 3′-end pattern (default sws ssw sww wss www) — restricts the last bases of each primer; use S=G/C, W=A/T, N=any. Multiple space-separated patterns are tried and the best result is kept.
  • Forward / Reverse tail (5′-3′) — optional universal tails appended to the 5′ end of each primer (e.g. M13 sequences) for downstream universal-primer detection.
  • Non-specific priming control and Low complexity priming control (both on by default) — see Common parameters → Mask repeats.
  • Allow overlapping primer positions — relaxes positional constraints when the search space is tight (use cautiously).
  • Multiplex PCR — designs assay sets for all input sequences to run together in one reaction. Each candidate primer/probe is checked against every primer/probe already picked for other sequences, so only mutually dimer-compatible sets are kept; sequences with few candidates are prioritized first so they aren't crowded out. Requires ≥2 input sequences (ignored for a single sequence). See Output tabs and columns for how the assay-set table changes when this is on.
  • C → T bisulfite conversion — switches the design into bisulfite-treated mode for methylation analysis (non-CpG cytosines are converted in silico; CpG sites preserved).

  • Input Sequences — your FASTA / retrieved sequences with live formatting and a per-tab counter.
  • My Primers — paste pre-existing primers/probes (optional) to be checked for compatibility with newly designed sets, e.g. for adding the new variant to an existing multiplex panel.
  • Report (Primer list) — flat list of every primer and probe candidate with their physico-chemical properties.
  • qPCR Assay Sets — final ready-to-order assays grouped by target: forward primer + reverse primer + allele-discriminating probe. With Multiplex PCR checked, this tab instead lists complete panels: each block groups one compatible assay per input sequence, and successive blocks are alternative panels — the number of blocks is capped by whichever sequence yielded the fewest compatible sets.

Per-oligo and per-pair columns follow the Common parameters → Standard output columns (ID · Sequence · Length · Tm · CG% · LC% · YR%, plus Fragment size and Annealing Tm on each pair).

  1. Load the variant — paste FASTA with [REF/ALT], upload a file, or retrieve by rsID from Ensembl. Use Load Example to see the expected formatting.
  2. Pick a probe format (TaqMan by default; switch to MGB for short / AT-rich targets, or Molecular beacon for difficult discrimination).
  3. Adjust Tm range, product size and 3′-end pattern if your master-mix or chemistry requires it.
  4. (Optional) Paste universal forward/reverse tails or pre-existing primers in My Primers.
  5. Click Generate. Review the assay sets in the qPCR Assay Sets tab.
  6. Copy the result area (Ctrl+ACtrl+C) and paste into Excel for ordering — see Exporting results.

  • No assay set returned — widen the Tm window / product size, or uncheck Non-specific priming control / Low complexity priming control; full checklist in Troubleshooting → Results are empty.
  • Probe not designed — for AT-rich variants switch to MGB; for very low-complexity flanks try Molecular beacon; check that the polymorphic position has at least 30 bp of unique flanking sequence on each side.
  • Ensembl retrieval fails — confirm the rsID and species id are correct, see Sequence retrieval fails.
  • Adding the assay to an existing panel — paste the existing primer/probe set into My Primers; new candidates incompatible with the panel will be filtered out.

castPCR — Competitive Allele-Specific TaqMan PCR

Applies to: castPCR Primer Design

castPCR (Competitive Allele-Specific TaqMan PCR) is a real-time PCR chemistry developed for picking out rare somatic mutations from a high background of wild-type DNA. The tool designs three components per locus that work as a competitive set in a single reaction: an allele-specific primer (ASP) with its 3′-end placed on the variant base, a common counter-direction primer shared across all alleles, and a common TaqMan or MGB detection probe placed in the conserved amplicon sequence between the ASP and the common primer. The probe detects all amplicons regardless of allele — allele discrimination comes entirely from the ASP. The chemistry routinely resolves mutant alleles down to ~0.1% mutant allele frequency — well below the limit of Sanger sequencing (~5–20%). A Multiplex castPCR option designs one dimer-compatible castPCR set per locus so that several targets can be amplified together in a single multiplex reaction — see Primer / probe design parameters and Output tabs and columns below.

Each well couples three components targeted at the same locus:

  • Allele-specific primer (ASP) — 3′-end placed on the variant base; mismatches at the 3′-end suppress extension on the non-target template, so only the matched allele amplifies efficiently. The tool generates one ASP per allele in both forward and reverse orientations.
  • Common TaqMan / MGB detection probe — placed in the conserved amplicon region between the ASP and the common reverse primer. The same probe detects amplicons produced by any ASP; it is shared across all alleles in the locus. Allele discrimination is determined entirely by the ASP, not by the probe.
  • Common counter-direction primer — shared across all alleles in the locus; works with every ASP to produce an amplicon of the same length. Selected so that the amplicon falls within the configured PCR product window for every allele.

The competitive design pushes the limit of detection to roughly 0.1 % mutant allele frequency — about one variant cell in a thousand wild-type cells.

Strengths

  • Sensitivity. Detects mutant fractions ≥0.1%, compared with the ~5–20% floor of Sanger sequencing.
  • Speed. Sample-to-result in under three hours on a standard real-time instrument.
  • Sample tolerance. Compatible with fresh-frozen tissue, cell lines, cfDNA, and FFPE material.
  • Robustness. Less affected by common PCR inhibitors (e.g. melanin in melanoma extracts) than several alternative allele-specific methods.

Typical applications

  • Targeted oncology. Genotyping actionable hotspots in EGFR, BRAF, KRAS and similar genes to inform therapy selection.
  • Resistance monitoring. Tracking minor resistant sub-clones (e.g. FLT3 in AML) that drive relapse.
  • Liquid biopsy. Detection of circulating tumor DNA and rare tumor cells from peripheral blood.

  • Paste FASTA in the Input Sequences tab with the polymorphism enclosed in square brackets (e.g. [A/G], [ATCG/-], [R]).
  • Upload a FASTA file — plain FASTA, multi-FASTA or sequences already containing [REF/ALT] brackets.
  • Retrieve from Ensembl by rsID — choose a species from the drop-down (Human, Mouse, Rat, Cow, Pig, Chicken, Arabidopsis), enter rsIDs separated by spaces, set the flank size (≥100 bp) and click Retrieve from Ensembl. The forward-strand context with the allele bracket is fetched automatically.
  • Custom species — type any Ensembl species id (lowercase, underscores, e.g. canis_lupus_familiaris) into the "Other species" field and click Add. Custom species are stored in your browser (localStorage) for later use.
  • Retrieve from NCBI by rsID — a second panel that goes to dbSNP instead of Ensembl, for human variants: enter one or more rsIDs in the Human rsIDs field, set the flank size and click Retrieve from NCBI. Use it when Ensembl is unavailable or the rsID is not in the selected Ensembl species. The retrieved records replace whatever is in the input box.

Provide ≥250 bp of flanking sequence on each side of the variant so the primer search has enough room.

  • Biallelic SNP: flank [A/G] flank — alleles separated by /
  • Shorthand (IUPAC): flank [R] flank — any IUPAC ambiguity code
  • Tri- / tetra-allelic SNP: [A/C/G], [A/C/G/T] or [N]
  • Insertion / deletion: [ATCG/-] (deletion) or [-/T] (insertion)
  • Excluded region — wrap with /.../ to keep primers and probes out; see Targeting markup.

All IUPAC ambiguity codes are accepted inside and outside the brackets — see Input formats → Allowed nucleotide codes.

Select the probe chemistry — one format at a time; checking one option clears the other. The probe is placed in the conserved amplicon region shared by all alleles (between the ASP and the common reverse primer) and detects all amplicons regardless of allele. Allele discrimination comes entirely from the ASP.

MGB (12–21 nt)

Minor Groove Binder conjugate. Shorter probes with higher Tm and better single-mismatch discrimination — preferred for AT-rich targets and tightly clustered SNPs. Default for castPCR.

TaqMan (18–30 nt)

Classic 5′-nuclease hydrolysis probe with a 5′ reporter dye and 3′ quencher. Use when SNP flanks are GC-balanced and a longer probe Tm (~+8 °C above primer Tm) is acceptable.

  • Length range (nt) — default 20–22. Defines the search window for the common counter-direction primer.
  • Tm range (°C) — default 61–63. Match it to your master-mix recommendation. Probe Tm is internally targeted ~5–10 °C above primer Tm (TaqMan: +8 °C, MGB: −2 °C of the primer Tm window).
  • PCR product size (bp) — default 100–200, already short enough for real-time chemistry.
  • SNP at 3′-end of ASP — number of bases at the 3′-end of the allele-specific primer that must differ between alleles (default 2; use 1 for simple SNPs and increase for tri-/tetra-allelic or short InDels). A larger value adds extra mismatches on non-target templates and suppresses leak amplification.
  • 3′-end pattern (default sws ssw sww wss www) — restricts the last bases of the common primer; S=G/C, W=A/T, N=any. Multiple space-separated patterns are tried and the best primer is kept.
  • ASP tail / LSP tail (5′-3′) — optional universal 5′ tags (e.g. M13 sequences) for downstream universal-primer detection. The first is appended to every allele-specific primer (F_ASP_…), the second to the common locus-specific primer (R_LSP_…).
  • Non-specific priming control and Low complexity priming control (both on by default) — see Common parameters → Mask repeats.
  • Allow overlapping primer positions — relaxes positional constraints when the search space is tight (use cautiously).
  • Multiplex castPCR — designs one dimer-compatible castPCR set per locus so that all input targets can be amplified together in a single reaction. Each candidate set for a locus is screened for primer-dimers against the primers and probes already committed for the other loci, so only mutually compatible sets are kept. Requires ≥2 input loci (ignored for a single sequence). See Output tabs and columns for how the assay-set tab changes when this is on.
  • C → T bisulfite conversion — switches the design into bisulfite-treated mode for methylation analysis (non-CpG cytosines are converted in silico; CpG sites preserved).

  • Input Sequences — your FASTA / retrieved sequences with live formatting and a per-tab counter.
  • My Primers — paste pre-existing primers/probes (optional) to be checked for cross-dimers against newly designed candidates, e.g. when adding a new variant to an existing panel.
  • Report (Primer list) — flat list of every candidate ASP, common detection probe and common primer with their physico-chemical properties.
  • castPCR Assay Sets — final ready-to-order assays grouped by target. For each locus two sets are reported: the forward set (ASP_F per allele + one shared Probe_F + R common primer) and the reverse set (ASP_R per allele + one shared Probe_R + F common primer). The probe is the same for all alleles within each orientation; pick the orientation that gives the better probe placement and allele discrimination. With Multiplex castPCR checked, this tab instead reports a single combined panel — one dimer-compatible castPCR set per locus (forward or reverse, whichever fits), chosen so the whole panel can be amplified together in one reaction. A locus for which no set compatible with the rest of the panel can be found is omitted.

Per-oligo and per-set columns follow the Common parameters → Standard output columns (with Fragment size and Annealing Tm on the common-primer row).

Cross-locus combinatorial use

All listed candidates are pre-checked for cross-dimers between sequences, so a chosen ASP + probe + common primer set from one locus can be combined with any other locus in a low-multiplex run without re-validation. To have the tool assemble those cross-compatible panels for you automatically, enable Multiplex castPCR (see above).

  1. Load the variant — paste FASTA with [REF/ALT], upload a file, or retrieve by rsID from Ensembl. Use Load Example to see the expected formatting.
  2. Pick a probe format (MGB by default; switch to TaqMan when the SNP flanks are GC-balanced and longer probes are acceptable).
  3. Adjust Tm range, product size and the SNP at 3′-end of ASP value for your variant type (1 for simple SNPs, 2+ for tri-/tetra-allelic or short InDels).
  4. (Optional) Paste universal forward/reverse tails or pre-existing oligos in My Primers.
  5. Click Generate. Review the assay sets in the castPCR Assay Sets tab — both forward and reverse triplets are reported; choose the orientation with the cleaner discrimination.
  6. Copy the result area (Ctrl+ACtrl+C) and paste into Excel for ordering — see Exporting results.

  • No assay set returned — widen the Tm window / product size, or uncheck Non-specific priming control / Low complexity priming control; full checklist in Troubleshooting → Results are empty.
  • Probe not designed — for AT-rich variants switch to MGB; check that the polymorphic position has at least 30 bp of unique flanking sequence on each side.
  • Poor allele discrimination at qPCR — increase SNP at 3′-end of ASP from 1 to 2 so the last two bases of the ASP differ between alleles; the additional mismatch on the non-target template suppresses leak amplification.
  • Ensembl retrieval fails — confirm the rsID and species id are correct; see Sequence retrieval fails.
  • Adding the assay to an existing panel — paste the existing primer/probe set into My Primers; new candidates incompatible with the panel are filtered out.

SNaPshot — SBE multiplex SNP genotyping

Applies to: SNaPshot Multiplex SNP Genotyping — SBE Primer Design

SNaPshot (Applied Biosystems) is a Single-Base Extension (SBE) chemistry for multiplex SNP genotyping. The locus is first amplified by conventional PCR; an SBE probe whose 3′-end anneals one nucleotide before the SNP is then extended by a single fluorescently-labelled ddNTP, identifying the allele by colour. Multiple loci are resolved on a capillary electrophoresis (CE) instrument by staggering probe lengths, so each locus appears as a distinct peak. The tool designs both the PCR primer pair and the SBE probe (in both forward and reverse orientations) for each input locus, and pre-validates the primers for cross-locus compatibility in a single multiplex PCR.

  • For each input locus: a PCR primer pair flanking the SNP and one or two SBE probes (forward and/or reverse orientation), so the more specific orientation can be chosen.
  • A panel-level SBE probe table with 5′-tails added to length-stagger the probes for CE separation.
  • A multiplex assay set in which every PCR pair is dimer-free against the primers of every other locus — pairs are interchangeable across loci.
  • Standard quality metrics for each oligo: length, Tm, CG%, linguistic complexity (LC%), purine/pyrimidine balance (YR%).

Paste a multi-FASTA where each sequence contains one SNP marked in square brackets. Provide ~150–300 nt of flanking context on each side so the tool has room to choose PCR primers and probes.

  • Biallelic SNP: ...flank[C/T]flank...
  • Multi-allelic: ...flank[G/A/C/T]flank...
  • IUPAC shorthand: ...flank[R]flank... — any IUPAC ambiguity code
  • InDels: ...flank[ATCG/-]flank... (one allele empty for a pure deletion)
  • Excluded region — wrap with /.../ to keep PCR primers and SBE probes out; see Targeting markup.
>rs3918290
gccacatacagtgaaaaccaactcaataa…ttagatgttaaatcacactta[C/A]gttgtctggaaagtcagcc…aatatacat
>rs1801133
tggggggaaaattagaggtaaccaaaatgg…aagaatgtgtcagcctcaaagaaaagctgcgtgatgatgaaatcg[Y]ctcccgcagacaccttctcc…ggaggtggctcagag

Sequences can also be retrieved by rsID directly from the input tab — from Ensembl (pick a species) or, for human variants, from NCBI/dbSNP with Retrieve from NCBI.

  • Length range (default 20–22 nt) and Tm range (default 61–63 °C). Tight Tm windows are essential for multiplex — all loci must amplify under one cycling profile.
  • PCR product size (default 200–500 bp). For SNaPshot it is the SBE probe — not the amplicon — that resolves the genotype, so the amplicon can be small.
  • 3′-end pattern (default sws ssw sww wss www): a space-separated list of 3-letter composition codes (S=G/C, W=A/T, N=any) describing the last three bases. The tool tries every pattern and keeps the best primer; the default favours one G/C clamp without G/C runs.
  • Forward / Reverse primer 5′-tails (e.g. M13 universal sequences): added to the 5′ end of every primer for compatibility with downstream sequencing or universal-tail PCR strategies. Leave empty if no tail is required.
  • Mask repeats & low-complexity regions — see Common parameters → Mask repeats.
  • Allow overlapping primer positions: returns multiple alternative primers covering the same window. Useful for finding workable candidates in difficult templates.

A pre-existing primer list (My Primers tab) is checked against every newly designed primer — any new primer that would form a heterodimer with one of yours is rejected.

An SBE probe has its 3′-end fixed one nucleotide before the SNP on the PCR amplicon. After the SAP/ExoI clean-up, the probe is extended by a single fluorescent ddNTP whose colour identifies the allele. The tool generates the probe in both orientations (forward and reverse) and you choose the one with the better local context (no nearby secondary SNPs, balanced Tm, no homopolymer at the 3′-end).

  • Probe core length 16–35 nt with the same Tm window as the PCR primers, so the SBE step works under the standard SNaPshot cycling profile.
  • Length staggering for CE: probes are sorted by core length and a 5′-tail is appended so adjacent probes differ by ≥6 nt in total length (comfortably clear of the CE peak resolution limit). The shortest probe gets no tail; the next gets just enough tail to be 6 nt longer; and so on.
  • Tail composition: by default the tail is poly-T. If you fill in Forward primer tail / Reverse primer tail, those sequences are also reused as the SBE-probe tail (cycled if shorter than required). Tails are printed in UPPERCASE to visually separate them from the lowercase probe core, e.g. TTTCACacgtacgtacgtacgtac.
  • Mobility caveat: tail composition affects CE mobility almost as much as length. Pure poly-dT is the safest baseline; mixed-composition tails (e.g. dGACT) shift mobility differently and may compress two adjacent probes into one peak. If using a non-poly-T tail, increase the minimum length gap or run a CE pre-test on the panel.
  • The tool does not account for the 1-nt extension and dye-induced mobility shift after SBE (typically 1–2 nt depending on dye/ddNTP). For very dense panels, add a 1–2 nt safety buffer to the staggering.

The full panel-level SBE probe table appears in the Multiplex Assay Sets tab under the heading === SBE probe panel ===.

Multiplex compatibility is enforced at design time, not as a post-hoc filter. As primers are designed locus-by-locus, every new candidate is checked against the cumulative pool of all previously accepted primers from every other locus. Any primer that would form a heterodimer with an already-accepted primer is rejected and an alternative is chosen.

  • For each locus you may see one or several PCR primer pairs in the Multiplex Assay Sets tab.
  • Any forward + reverse pair from one locus can be freely combined with any pair from any other locus in a single mPCR. All listed pairs are pre-validated to be cross-locus dimer-free.
  • This lets you choose, per locus, the pair with the best Tm match, smallest amplicon, or highest specificity, without re-checking compatibility against the rest of the panel.

Note: cross-locus compatibility is enforced for the PCR primers. The SBE step uses different chemistry and is run after SAP/ExoI removes excess PCR primers and dNTPs, so probe-vs-PCR-primer interactions are not the limiting factor in practice.

Two output tabs serve different purposes:

  • Report (Primer list) — every candidate primer and probe found per locus, with full quality metrics. Use this to inspect alternatives or to debug a locus that produced no multiplex set.
  • Multiplex Assay Sets — only the validated PCR primer pairs (cross-locus dimer-free) plus the panel-level SBE probe table with length-staggering tails. This is the orderable output.

Column headers in both tabs are tab-separated and Excel-friendly:

ID    Sequence(5'-3')    Length    Tm(°C)    CG(%)    LC(%)    YR(%)    Fragment_Size(bp)/Tm(°C)
  • LC% (linguistic complexity) penalises repeats and homopolymers — values ≥70% are recommended for primers.
  • YR% reports purine/pyrimidine balance — extremes can predispose to secondary structure.
  • Fragment size / Annealing Tm is shown on the reverse-primer row of each PCR pair: the predicted amplicon size in bp and a Tm estimate of the full amplicon (informational, not the cycling Tm).
  • Probe ID encodes locus and orientation, e.g. 1F_… = locus 1, forward probe.

When C→T bisulfite conversion is enabled, the input sequence is first virtually converted (every unmethylated C becomes T on each strand independently), and primers/probes are designed against the converted sequence. This is the standard preparation for bisulfite-sequencing assays where the methylation status at a CpG site is read out as a C/T polymorphism after conversion.

  • Mark the CpG of interest with [C/T] in the original (unconverted) sequence — this represents methylated vs unmethylated allele after conversion.
  • Avoid placing PCR primers over additional CpG sites whenever possible (the tool does not do this automatically — use /.../ markup if needed).
  • The two strands diverge after conversion, so the forward and reverse SBE probes are designed independently on each converted strand.

  1. Enter sequence(s) — paste a multi-FASTA with each SNP marked as [REF/ALT] (e.g. …flank[C/T]flank…), upload a local FASTA, or retrieve flanking context by rsID from Ensembl.
  2. Set parameters — PCR primer length / Tm / amplicon size, optional 3′-end composition pattern, and 5′ mobility-shift tails for the forward and reverse SBE primers (poly-dT or poly-dGACT). Tick Mask repeats to avoid low-complexity flanks; tick C→T conversion for bisulfite-methylation assays.
  3. Press Generate — the Report tab lists every candidate PCR primer and SBE probe per locus; the Multiplex Assay Sets tab returns validated, dimer-free PCR + SBE combinations ready for ordering. Pairs from different sequences are pre-validated to be cross-locus compatible, so any F+R pair from one locus can be mixed with any pair from any other locus in the same mPCR.
  4. Plan the multiplex — stagger SBE primer lengths by ≥6 nt (using the 5′ tails) so all loci resolve as distinct peaks on capillary electrophoresis.

  1. Multiplex PCR — pool all forward and reverse primers from the chosen pairs at equal final concentration (typically 0.05–0.2 µM each). Use a hot-start polymerase and the Tm window reported by the tool as the annealing temperature.
  2. SAP/ExoI clean-up — remove unincorporated dNTPs and excess primers (Shrimp Alkaline Phosphatase + Exonuclease I). This is mandatory before the SBE step.
  3. SBE reaction — combine the cleaned PCR product with the pooled SBE probes and the SNaPshot reaction mix (contains the four fluorescent ddNTPs and a thermostable polymerase). Cycle per the manufacturer's protocol.
  4. Post-extension SAP — dephosphorylate unincorporated fluorescent ddNTPs (otherwise they appear as dye blobs on the trace).
  5. Capillary electrophoresis — run with a size standard (e.g. GeneScan 120 LIZ) and analyse with GeneMapper or comparable software. Each locus appears as a coloured peak at the expected probe length.

If two adjacent loci share a peak position in the trace, increase the length-stagger gap or swap one probe to the opposite orientation.


TSDR Probe Designer

Applies to: TSDR Probes — toehold exchange probes for SNP/allele discrimination

Designs a two-strand toehold exchange probe (signal strand C + Protector P) for a target you supply with the variant site marked as [A/G] (or an IUPAC code such as [R]). The design is tuned so the strand-displacement reaction is near-thermoneutral for the on-target allele, which maximises single-base discrimination. Method: Zhang, Chen & Yin (Nat Chem 2012); kinetics scale with toehold ΔG per Zhang & Winfree (JACS 2009).

The target allele X reacts with the resting probe by toehold exchange: X + C·P ⇌ X·C + P, with ΔG°rxn = ΔG°(X·C) − ΔG°(C·P).

  • On the target the probe footprint is a forward toehold F (single-stranded on C) plus a branch-migration domain M that carries the SNP.
  • The signal strand adds a synthetic reverse toehold (clamp) R that only the Protector complements — the target cannot invade it. Its length/sequence is tuned automatically so ΔG°rxn ≈ target (default 0).
  • At thermoneutrality a single mismatch in M shifts ΔG°(X·C) by ΔΔG>0, driving the off-target reaction uphill. The discrimination factor Q (on-yield ÷ off-yield) is largest here.

Each probe is complementary to one allele. Tick Two-colour mode to design one probe per allele in a single run: the output then covers the whole assay — a reporter dye per probe, per-probe thermodynamics, a cross-probe hybridisation scan and a dye spectral-compatibility check.

  • Paste one target sequence (5′→3′), with or without a FASTA header.
  • Mark the variant as [A/G] (slash-separated alleles) or a single IUPAC code such as [R], [Y], [S].
  • The input is checked before the design runs. A missing or unbalanced bracket, an empty [], protein input, or fewer than 12 bases refuses the run: nothing is designed, the reason appears in the message strip above Design probe, and the result pane says the input was refused. Warnings — characters that will be ignored, for example — are shown but do not block. Only the first record is examined and used.
  • Substitutions — single-base, or multi-base of equal length such as [AT/GC] — short indels ([/ATG] insertion, [ATG/] deletion) and a fixed site with no polymorphism ([A]) are all designed for. The marked segment must be 12 nt or shorter or the input is refused, and beyond about 6 nt an indel becomes a presence/absence assay rather than a thermoneutral discrimination.
  • At least 12 bases are required, or the run is refused — and that is only a floor. Provide enough flanking sequence on both sides to fill the footprint (default 24 nt) so the variant can sit centrally in the branch domain; with shorter flanks the footprint is clamped to what is there, the variant sits off-centre, and the Warnings tab says so.
  • Neither opening the page nor loading the example computes anything — press Design probe to run the design. Changing any parameter afterwards recomputes it in place, and will run the design even if the button has not been pressed yet.

  • Probe (C) / Target (X) concentration (µM): set the equilibrium yields and the reported Q.
  • Operating temperature (°C): ΔG is computed at this temperature from ΔH/ΔS.
  • Total salt / Mg²⁺: entropy salt correction (monovalent-equivalent).
  • Probe footprint (nt) and toehold length range (nt): the search space for F and M.
  • Two-colour mode: designs one near-thermoneutral probe per allele in a single run, instead of one probe for the first allele. Ignored for a monomorphic [A] target, which has only one allele.
  • Target ΔG°rxn (kcal/mol): 0 = thermoneutral = maximal discrimination. A monomorphic [A] detection probe has no allele to discriminate and reports no Q — set a negative value there (e.g. −4) so the target strongly displaces the protector; thermoneutral gives only about 50% yield.

  • Probe design: the F/M/R segments, the final C and P strands (5′→3′), a schematic, and the discrimination factor Q.
  • Thermodynamics: per-allele ΔG(X·C), ΔG(C·P), ΔG°rxn, ΔΔG and equilibrium yield χ; the forward/reverse toehold ΔG; and the forward-toehold length scan.
  • Warnings: SNP placement, toehold-window flags, low Q, self-structure/leak, and the modelling caveats below.
  • In Two-colour mode the same three tabs describe the whole assay rather than one probe: a dye table (reporter, emission maximum, detection channel, on-allele and Q) followed by a labelled block per probe; per-probe thermodynamics; and, in Warnings, a cross-hybridisation scan between the signal/protector strands of different probes plus a dye spectral-compatibility check.

  • Equilibrium, not kinetics. Yields and Q are equilibrium values; forward/reverse toehold ΔG are reported as rate proxies (Zhang & Winfree 2009), not a simulation.
  • Mismatch thermodynamics are measured. Off-target alleles use internal single-mismatch nearest-neighbour parameters (Allawi & SantaLucia 1997/1998; Peyret et al. 1999; SantaLucia & Hicks 2004), not step-zeroing.
  • Salt is approximate. A monovalent-equivalent entropy correction (SantaLucia 1998; von Ahsen 2001 Mg term) is applied; ΔG°rxn is a difference of similar duplexes, so salt largely cancels.
  • The clamp is synthetic sequence. Confirm it is free of secondary structure and cross-hybridisation in your assay context.

MLPA / digitalMLPA — Multiplex Ligation-dependent Probe Amplification

Applies to: MLPA Synthetic Probe Design

MLPA measures the copy number of up to 50 genomic targets in a single tube. Each target is interrogated by a probe of two half-oligonucleotides that hybridise to immediately adjacent sites; a thermostable ligase joins them only when the ligation junction is a perfect match. A single universal primer pair then amplifies every ligated probe, and because each probe carries a length-tuning stuffer, the products separate by size on capillary electrophoresis (CE) the peak area of each probe is proportional to the copy number of its target. Unlike PCR-based assays, it is the probe, not the genomic DNA, that is amplified. This tool designs both half-probes for every input target and assembles a length-staggered, dimer-checked panel. The same hybridising arms also feed an NGS readout digitalMLPA where each probe is identified by sequence rather than amplicon size, scaling a single reaction to ~1000 targets: swap the CE universal-primer tags for Illumina adapters plus a per-sample barcode oligo (the length-staggering stuffers then become optional).

  • For each input target: a Left Probe Oligo (LPO) and a Right Probe Oligo (RPO). With a [REF/ALT] marker the junction is placed at that site — one LPO per allele, common RPO; otherwise the best repeat-free, G/C-clamped site is chosen.
  • The two target-specific hybridising arms (LHS, RHS) with length, Tm, GC% and complexity, plus their position on the input strand.
  • A panel-level table where every probe is sized to a unique amplicon length via a per-probe stuffer, ready for one-tube CE.
  • The universal primer pair shared by the whole panel and a cross-dimer check over all arms and primers.

A synthetic MLPA probe is two oligos that, once ligated, read 5′→3′ as a single PCR template:

LPO (5'->3')  =  [forward tag] + LHS
RPO (5'->3')  =  5'P + RHS + [stuffer] + [reverse tag]

amplicon      =  [forward tag] + LHS + RHS + [stuffer] + [reverse tag]
  • LHS / RHS — the target-specific left and right hybridising sequences. Concatenated, they equal a contiguous window of the input strand, so the ligation junction is shown directly in the input, between LHS and RHS; the half-probes anneal to the complementary strand. The LHS provides the 3′-OH and the RHS the 5′-phosphate at the nick.
  • Forward tag (LPO 5′-end, default GGGTTCCCTAAGGGTTGGA) and reverse tag (RPO 3′-end, default TCTAGATTGGATCTTGCTGGCAC) — the universal primer sequences. The reverse PCR primer is the reverse complement of the reverse tag (GTGCCAGCAAGATCCAATCTAGA). Both are editable; the defaults are the standard SALSA MLPA primers.
  • Stuffer — filler added to the RPO so each amplicon reaches a unique length. Printed in UPPERCASE to separate it from the lowercase target arms. The stuffer is tiled from the source pattern and is not automatically screened for secondary structure or complexity, so inspect a custom or long stuffer (and avoid homopolymers/strong hairpins).

The RPO must be ordered 5′-phosphorylated; without it the ligase cannot seal the nick.

Paste a multi-FASTA with one target region per record — typically an exon or dosage-sensitive segment of ~120–300 nt. The optional [REF/ALT] markup decides where the ligation junction goes.

  • Allele-specific (1–4 alleles) — mark the interrogated site and the junction is placed at it, with the allele as the 3′-end of the LHS: a separate LPO is emitted per allele while the RPO is common. Works for SNPs ...gtcaccat[C/T]tatggg... (or the IUPAC ambiguity code [Y] = C/T, [R] = A/G, [N] = all four, …), InDels [CCC/T], and deletions [CCC/] or [CCC/-] (the empty side is the del allele, whose LHS is just the upstream flank). Only the perfectly matched allele ligates. Because alleles can differ in length, their amplicons can differ in size — so InDel alleles resolve by size on CE, while equal-length SNP alleles share one size (tell them apart by ligation/sequence).
  • Position marker — a single-token or empty bracket ([G], []) just fixes the junction location; one probe is designed there.
  • No bracket — for plain copy-number targets the best repeat-free, G/C-clamped ligation site is chosen automatically.
  • Targets can also be retrieved by rsID directly from the input tab — from Ensembl (pick a species) or, for human variants, from NCBI/dbSNP with Retrieve from NCBI; the returned variant is bracketed automatically and becomes the ligation site.

The optional Existing Oligos tab lets you supply oligos that the new probe arms should not dimerise with.

  • Hybridising arm length (default 21–30 nt). Each arm grows to the shortest length that reaches the minimum Tm, within this window.
  • Min Tm per arm (default 64 °C, nearest-neighbour). Hybridisation/ligation runs ~60 °C overnight, so each arm must melt comfortably above it; MRC-Holland recommends arms several °C above the hybridisation temperature.
  • Complexity — a linguistic-complexity floor (LC% ≥ 70) is applied internally to reject low-complexity arms. GC% is reported for each arm but is not constrained; Tm is the governing criterion.
  • G/C base pair at the ligation junction (default on) — the 3′ base of the LHS and the 5′ base of the RHS are constrained to G/C to stabilise the nicked junction. The LHS 3′ base is the key ligation-fidelity position; the RHS 5′ base is clamped as well as a conservative default. If no compliant junction exists, the tool relaxes this constraint and flags it.
  • G/C clamp at both outer arm ends (off by default) — additionally constrains the 5′-end of the LHS and the 3′-end of the RHS.
  • Mask repeats — exclude repeat-masked positions from the arms. A marked [REF/ALT] site is the design target (not avoided); any other brackets in the same record are kept off the arms.
  • Universal tags & stuffer source — editable; defaults are the standard SALSA primers and a balanced filler.

If a target is too short, too AT-rich to reach the Tm, or fully repeat-masked, that target is reported as skipped / no-probe rather than forcing a poor design.

Because all probes share one primer pair, they are resolved only by amplicon length. Probes are sorted by their stuffer-free length and a per-probe stuffer is added so each successive amplicon is at least the minimum size spacing longer than the previous one, starting at the minimum amplicon size.

  • Amplicon size range (default 130–500 nt) and min size spacing (default 8 nt). Wider spacing is safer; CE peak resolution and length-dependent mobility both degrade as peaks crowd. The 130 nt floor keeps the panel clear of the <120 nt control-fragment band (see note below).
  • Amplicon length = forward tag + LHS + RHS + stuffer + reverse tag. With the default tags that is 42 nt + arms + stuffer.
  • If the staggered sizes exceed the ceiling, the tool warns; reduce the probe count, shorten arms, or raise the maximum.

Commercial MLPA kits add separate denaturation/quantity control fragments below ~120 nt; leave headroom for them when choosing the minimum size.

Two tab-separated, Excel-friendly outputs:

  • Probe Report — per target: the ligation junction position, the LHS and RHS with length/Tm/GC/LC% and coordinates, then the full LPO and RPO sequences to synthesise (tags UPPERCASE, target arms lowercase). For an allele-specific site, one LHS/LPO line is listed per allele above the common RHS/RPO; allele probe IDs are suffixed with the allele — [C], [T], [CCC], [del], …
  • Probe Panel — the universal primer pair, a size-sorted table (ProbeID, Locus, Amplicon, LHS, RHS, Stuffer, Ligation@, TmL/TmR), the full ligated probe sequences, and the cross-dimer summary.

Ligation@ is written as a^b — the junction sits between input nucleotides a (LHS 3′-end) and b (RHS 5′-end); for an allele-specific site it reads p(C), the variant position and allele. Equal-length alleles (SNPs) share one amplicon size — distinguish them by allele-specific ligation/sequence (digitalMLPA); length-varying alleles (InDels) resolve by size on CE.

  1. Enter targets — paste one region per FASTA record (exon / dosage segment), upload a FASTA, or fetch by rsID from Ensembl. Mark any known variant to avoid with [REF/ALT].
  2. Set parameters — arm length / Tm / GC, the amplicon size range and spacing, and (if needed) custom universal tags or stuffer. Keep the junction G/C clamp on for ligase fidelity.
  3. Press Design Probes — review the Probe Report per target, then order the LPO (unmodified) and RPO (5′-phosphorylated) oligos from the Probe Panel.
  4. Run MLPA — denature DNA, hybridise the probe mix overnight at ~60 °C, ligate, then PCR with the single universal primer pair (label the forward primer, e.g. 6-FAM) and size the products on a CE instrument. Normalise each peak area to reference probes to call copy number.

Always include reference (copy-number-stable) probes and run several normal controls; MLPA reports relative dosage, so calls depend on the reference set and on consistent DNA quality.

  • Jan P. Schouten, Cathal J. McElgunn, Raymond Waaijer, Danny Zwijnenburg, Filip Diepvens, Gerard Pals Relative quantification of 40 nucleic acid sequences by multiplex ligation-dependent probe amplification. Nucleic Acids Research 2002, 30(12): e57. doi:10.1093/nar/gnf056
  • Principle of MLPA

STR / SSR — microsatellite fragment-analysis primer design

Applies to: STR / SSR Microsatellite Genotyping

A microsatellite (STR / SSR) is a tandem repeat of a short motif (typically 2–6 nt — di- to hexanucleotide) whose copy number varies between alleles. Flanking PCR primers turn each allele into a product of a distinct length, sized on a capillary-electrophoresis (CE) instrument against an internal size standard. The tool locates the repeat in each input, designs at least three and up to twelve Tm-matched, dimer-checked alternative primer pairs in the unique flanks (each with a different forward and/or reverse primer), adds economical universal fluorescent labelling (the M13-tail / Schuelke method), and assembles a multiplex panel — assigning each locus a dye channel and, if needed, a different alternative pair — so that product sizes stay a minimum gap apart within any one colour (or across the whole panel for a single-dye layout).

  • For each input: the detected microsatellite (motif, copy number, position) and its alternative flanking primer pairs — at least three, and up to twelve when many loci must share one dye channel — ranked by amplicon size and Tm match, each with length, Tm, GC%, linguistic complexity (LC%) and purine/pyrimidine balance (YR%) for both primers.
  • The M13-tailed forward and PIG-tailed reverse sequences to order, plus the expected product size and allele size-range.
  • A multiplex panel assigning each locus a dye channel + size window (kept a minimum gap apart within one colour — or across the whole panel for a single-dye layout), the universal labelled primer, and a cross-dimer audit.

Paste a multi-FASTA where each record contains one microsatellite with ≥ ~25 nt of unique flank on each side (e.g. paste the region reported by Repeats).

  • Auto-detect — the strongest tandem repeat (motif 2–6 nt, ≥ the minimum repeat length) is found, extended to whole copies and normalised to its primitive motif, so (CACACA) is reported as (CA); homopolymer (1 nt) runs are down-weighted as poor markers.
  • Marked — wrap the repeat in [..] (e.g. ...acctg[ca]n... or the literal repeat [cacacacaca…]) to fix it explicitly; everything inside the brackets is taken as the repeat.
  • Mononucleotide runs and very short flanks may yield no primer pair — widen the size range or supply longer flanks.

Direct dye-labelling of a primer per locus is expensive. The Schuelke (2000) method adds a universal 5′-tail (default M13(-21) TGTAAAACGACGGCCAGT) to the forward primer; a single dye-labelled universal primer equal to the tail is added to the PCR and labels every locus.

  • Run the tailed forward primer at ~¼ the reverse concentration plus the labelled universal primer; a two-step anneal (high then low) lets the tail prime in later cycles.
  • The reverse primer gets a 5′-PIG-tail (default GTTTCTT) that promotes full non-template +A addition by Taq, sharpening called peaks and avoiding split (−A/+A) alleles.
  • Both the M13 tail and the reverse PIG-tail add their length to every product — the reported sizes already include them (PIG-driven +A adds ~1 nt more).
  • For multi-colour panels, give each dye channel its own universal tail (e.g. the four tails of Blacket et al. 2012) by editing the tail field.

  • Primer length / Tm (default 18–27 nt, 57–63 °C) — a tight Tm window is important for multiplex co-amplification.
  • Product size (core, default 90–400 bp) — every Tm-valid combination in range is ranked by amplicon size and Tm match, and a size-spread selection that always includes the best-ranked pair is reported as the alternative pairs — at least three per locus, rising towards twelve as more loci have to share one dye channel; the multiplex panel picks whichever one best avoids a same-dye size collision (not always the best-ranked one). The lengths of both 5′-tails (M13 forward + PIG reverse) are added on top.
  • Buffer from repeat — keep the primer 3′-end this many nt away from the repeat edge (avoids slippage).
  • Allele range (± repeats) — defines each marker's expected size window (= ± repeats × motif length) for multiplex packing.
  • Dye channels — different dyes may share a size range freely (colour tells them apart); markers are packed greedily by size within each channel. Set this to 1 for a single-colour panel, where every product must instead be told apart purely by length.
  • Min size gap, one dye (nt) (default 15) — the minimum separation required between product sizes within the same dye channel (10–20 nt is typical for reliable CE resolution). The packer tries each locus's alternative primer pairs to fit; if none clear the gap on any channel, the locus is placed at its least-bad size and flagged ⚠.
  • Motif length (nt) (default 2–6, di- to hexanucleotide; the minimum cannot go below 2) — the window scanned by auto-detect; homopolymer (1 nt) runs are down-weighted, as they make poor STR markers.
  • Min repeat length (nt) (default 12) — the minimum tandem-repeat tract length to call a microsatellite (also enforced as ≥ 2× the motif length and ≥ 2 copies).
  • Non-specific priming control (default on) — keeps both primer arms off dispersed/interspersed repeats elsewhere in the flanking sequence (the microsatellite itself is already excluded by the buffer setting).
  • Low complexity priming control (default on) — keeps both primer arms off other SSR/microsatellite and telomere-like stretches in the flanks, the same blocks detected by Repeat Finder. Disable either control only if it prevents a primer from being found and the flanks cannot be replaced.

Always co-load an internal size standard (e.g. GeneScan LIZ) and bin alleles against a reference ladder; report repeat-unit allele names rather than raw bp where a nomenclature exists.


In silico PCR

Applies to: In silico PCR

  • Searches for primer/probe (or gRNA/miRNA-like) binding sites across provided sequences.
  • Predicts likely amplicons within configured product-size constraints.
  • Reports mismatches and provides a log of hits/off-targets.

  1. Paste/upload target sequences in FASTA.
  2. Provide the primer/probe list in the dedicated tab.
  3. Press Analysis and review both output tabs: Results and PCR amplicons.
  4. If no output: confirm hits exist within the allowed product size range, and verify primer orientation (forward/reverse).

LAMP primer design

Applies to: LAMP primer sets design tool

  1. Paste/upload target sequence in FASTA or retrieve by NCBI accession.
  2. Set primer constraints (length, Tm) and maximum F2–B2 amplicon size.
  3. Enable Loop Primer Design for faster amplification (LF/LB).
  4. Click Generate Primers and review both output tabs: Report (Primer list) and LAMP Primer Assays.

Core primer distances
  • F2 to B2 (5′ ends): set by the Max F2-B2 amplicon size field — default 200 bp, adjustable 100–350 bp. Only the upper bound is enforced; there is no minimum separation.
  • F2 to F3: 0–20 bp
  • B2 to B3: 0–20 bp
Loop forming regions
  • 5′ of F2 to 3′ of F1: 0–40 bp
  • 5′ of B2 to 3′ of B1: 0–40 bp

Primer typeTarget Tm
F1c / B1c / LF / LB (inner primers & loops) 65–67°C
F2 / B2 / F3 / B3 (outer primers) 60–62°C

The Tm range (°C) field on the tool sets the outer-primer window (default 60–62 °C); the inner and loop primers are always designed for that window plus 5 °C, so both rows move together when you change it. These Tm values are computed under the same fixed, standard-PCR-like conditions used by every design tool on this site (see Common parameters → Tm calculation conditions), not LAMP's typically higher Mg²⁺ and primer concentrations. Keeping one condition set makes Tm comparable across tools, but it will read a few degrees lower than a LAMP-specific calculator would report for the same primers — treat it as a relative design target, not the literal isothermal reaction temperature.

  • Use [ ... ] to constrain primer design to a specific region.
  • Exclude regions with /.../ (e.g. repeats); see Targeting markup.

Enabling C→T bisulfite conversion automatically adjusts the default length (17–28 nt), Tm (57–59°C), and overlapping primer settings — suitable for designing LAMP primers on bisulfite-converted templates for methylation analysis.


CRISPR guide RNA design

Applies to: CRISPR Guide RNA Designer

Scans a target for the PAM sites of the chosen nuclease, scores every protospacer, and searches the sequences you supply for off-targets. Mark the change you want to make and the same run also produces the HDR donor, the primers that screen the edit, a prime-editing pegRNA and the base-editing window. Off-target scoring follows Hsu et al. (Nat Biotechnol 2013); prime editing follows Anzalone et al. (Nature 2019); donor geometry follows Richardson et al. (Nat Biotechnol 2016).

Press Design guides once and five reports are filled in together, because they are five views of the same design decision:

  • Guides — every protospacer, ranked; then, per guide, the full sgRNA or crRNA and the cloning oligos for a BbsI-cut U6 vector.
  • Off-targets — each shown guide against every sequence you pasted, both strands, with the mismatch pattern and an MIT score.
  • Donor & primers — the ssODN with the edit and a PAM-blocking change, plus a screening primer pair around the cut.
  • Prime & base editing — the pegRNA with its PBS and RT template and the PE3/PE3b nicking guides; then the base-editing window with its bystanders.
  • Warnings — everything the design is uncomfortable about, and the modelling caveats behind every number.

Nothing runs when the page opens or when an example is loaded: a design belongs to the sequence it was run on, so editing the input retires the previous result rather than silently keeping it on screen.

  • Paste one target sequence, 5′→3′, with or without a FASTA header. Only the first record is used.
  • The first allele in the bracket is what the genome has; the second is what you install. [C/T] installs T where the genome reads C, [/ATG] inserts ATG, [ATG/] deletes it. [C] marks a site without an edit — useful for a knock-out, where the guides matter and the donor does not. [Y] and the other IUPAC codes expand to their two bases.
  • Leave the brackets out entirely and the tool tiles guides across the whole sequence, ranked by score alone. The donor, pegRNA and base-editing reports then have nothing to work with and say so.
  • Only one bracket is allowed. Two would leave it ambiguous which one the donor is built around, so the run is refused.
  • Paste 600–1000 nt around the site. The homology arms and the screening primers both need flanking sequence; 30 nt is the floor at which anything runs at all, not a useful working length.
  • The Off-target sequences tab takes as many FASTA records as you like — paralogues, pseudogenes, the donor plasmid, the vector backbone. The target itself is always searched too.

Two worked examples ship with the tool. Load Example is a β-globin coding region with an A→T substitution to install — a transversion, which exercises the donor and the pegRNA and shows why no base editor can make it. Example: base editing is a coding fragment where C→T turns a CAG codon into a stop, with one bystander cytosine inside the window; it arrives with the editor and the reading frame already set.

Eleven presets, from SpCas9 (NGG) through the relaxed-PAM variants (SpCas9-NG, SpRY, VQR, VRER), the compact orthologues (SaCas9, SaCas9-KKH, Nme2Cas9) to the Cas12a family (AsCas12a/LbCas12a on TTTV, enAsCas12a on TTYN), plus a Custom PAM entry where you set the PAM, which side of the protospacer it sits on, the spacer length and the cut offset.

  • Cas9 family — PAM 3′ of the protospacer, blunt cut 3 bp from it. The cut column is the forward coordinate the blade falls after.
  • Cas12a family — PAM 5′ of the protospacer, staggered cut roughly 18/23 nt from the PAM leaving a 5-nt 5′ overhang, so the cut column shows two numbers.
  • A relaxed PAM (NAG for SpCas9) is used in the off-target search only, scored at 30% weight. It never proposes a new guide: designing against a PAM the enzyme cuts poorly is a choice, not a default.
  • Sites whose protospacer or PAM contains an ambiguous base (N, R, Y…) are skipped and counted in the header line.

Cloning oligos follow the Zhang-lab convention for a BbsI-cut U6 vector: CACC + spacer, or CACCG + spacer where the spacer does not already start with G, because U6 initiates on a G. The added base is flagged in the output.

score — a transparent composite activity heuristic, 0–100, starting from a baseline of 90. Every term is a published sequence rule, and each guide's output prints the arithmetic:

  • GC outside the preferred window is penalised; GC in the 45–60% band earns a small bonus.
  • TTTT costs 35 points: it terminates Pol III, so a U6-driven guide is truncated.
  • Homopolymers of 5 or more, low linguistic complexity, an over-stable or too-weak seed, a T at the PAM-proximal position, a U-rich 3′ end, a hairpin inside the spacer, and pairing between the spacer 3′ end and the scaffold lower stem all subtract; a G at the PAM-proximal position and a spacer that already begins with G add.

It is not Rule Set 2 / Azimuth (Doench et al. 2016). That model is a gradient-boosted regressor with thousands of fitted parameters and cannot be reproduced honestly inside a small client-side script. Read this number as a filter that removes obviously poor guides, not as a predicted indel frequency, and break ties with the off-target table rather than with the last point of score.

spec — the MIT / Hsu (2013) aggregate: 100 ÷ (100 + Σ hit scores), so 100 means nothing was found and a low value means a strong competing site exists. Each hit is scored from its mismatch positions using the published position-weight vector. The weights were measured for a 20-nt SpCas9 spacer; for other spacer lengths the positions are mapped onto that scale and the report says the number is approximate.

This is not a genome-wide search. DigitalGens runs entirely in your browser: no genome is downloaded and no sequence leaves the machine, so the search covers exactly the sequences in the two input boxes and nothing else. A clean result means "clean in what you supplied" — never "clean in the genome". For a genome-wide count, run the chosen guides through a server-side index such as CRISPOR or Cas-OFFinder as well.

Within that scope the search is exhaustive: every PAM position on both strands of every record, compared against every shown guide up to the mismatch limit. What it is good at is the question a genome index answers badly — does this guide also cut the paralogue, the pseudogene, the donor plasmid or the second copy in my construct? Paste those in and the answer is exact.

  • Mismatch positions are counted from the PAM-distal end of the protospacer, so position 1 is the 5′ end for Cas9.
  • A hit whose mismatches all fall outside the seed is flagged seed intact — high risk: the PAM-proximal bases still pair, and such sites are cleaved efficiently however many distal mismatches there are.
  • A perfect match elsewhere in the supplied sequences is labelled as such and collapses the specificity score, which is the intended behaviour.

  • The donor is built around the top-ranked guide and spans both the cut and the edit, with the requested homology arm beyond each. The edit and any PAM-blocking change are in UPPER CASE; homology is lower case. Both strands are printed.
  • Strand. On Auto the ssODN is complementary to the PAM-carrying strand — the one Cas9 releases first, and the orientation Richardson et al. found most efficient. The Asymmetric donor option applies their 36 nt PAM-distal / 91 nt PAM-proximal geometry instead of symmetric arms.
  • Re-cut protection. A repaired allele that still matches the guide is simply cut again. The tool first checks whether the edit itself has already destroyed the PAM or put two mismatches in the seed — often it has, and then it says so and changes nothing further. Otherwise it looks for a single change that blocks re-cutting, trying the constrained PAM bases first and then the seed from the PAM inwards.
  • Silent changes need a reading frame. Set Coding strand and First base of the CDS and the search prefers a synonymous substitution and reports it as, say, N58N (silent). Without a frame it still protects the PAM but labels the change NOT verified as silent — which is a warning to check, not a formality.
  • Screening primers flank the cut with the amplicon deliberately off-centre (about 35%), so a T7E1/Surveyor digest gives two clearly different fragments. The report gives both fragment sizes, the annealing temperature, and the distance from the forward primer to the cut, which is what decides whether a Sanger trace is clean enough for TIDE/ICE.

Tm is the site-wide nearest-neighbour value (55 mM Na⁺/K⁺, 1 mM Mg²⁺, 0.2 µM oligo), so it is directly comparable with every other tool here. Run the pair through In silico PCR before ordering.

The nickase cuts the PAM-carrying strand 3 bp from the PAM; the freed 3′ end anneals to the primer-binding site (PBS) and is extended over the RT template (RTT), writing the edit into the genome. The edit therefore has to lie 3′ of the nick on that strand, and close to it.

  • PBS length is chosen by Tm. Every length in the range is scored and the one whose nearest-neighbour Tm lands closest to the target — 30 °C by default, following the prime-editing protocols — is used. GC-rich sites therefore get a shorter PBS and AT-rich sites a longer one, automatically. The report prints the Tm actually achieved.
  • RTT covers the nick-to-edit gap, the edit itself and the requested 3′ homology. If its 5′-most base would be a C the template is extended, since that base reduces efficiency (Anzalone 2019).
  • PE3 nicking guides nick the opposite strand 40–90 nt away. PE3b guides have a protospacer that spans the edit, so they only nick the edited allele — later in the process, and with far fewer indels.
  • Optional tevopreQ1 3′ motif (an epegRNA, Nelson et al. 2022) and the flip+extension scaffold.

The PBS anneals RNA to DNA; its Tm is computed with the DNA/DNA nearest-neighbour model, which runs a few degrees low for a hybrid. The 30 °C working target is defined on the same scale, so the comparison is self-consistent — but bracket the chosen PBS with two or three lengths either side experimentally, because the optimum is locus-specific and cheap to test.

A base editor deaminates every susceptible base inside a window of the protospacer, not only the one you meant. Six editors are offered — BE4max, the narrowed YE1, the broad evoAPOBEC1, ABE7.10, ABE8e and a C→G CGBE — each with its own window, counted from the PAM-distal end of the protospacer.

  • Guides are found where the target base falls inside the window, on either strand: a C→T on the minus strand is written G→A on the plus strand, and the tool handles that itself.
  • rel.act combines the position weight inside the window with the deaminase motif preference (TC ≫ GC for APOBEC1; TA > GA for TadA). It ranks candidates against one another and is not a predicted editing percentage.
  • Every other susceptible base in the window is listed as a bystander, with its genomic position, its motif, and — when a reading frame is given — the protein consequence of editing it. A silent bystander is a non-event; a missense or nonsense one may cost you the experiment.
  • A diagram marks the target (^), the bystanders (*) and the window (-) under the protospacer.

Base editors install transitions — C→T and A→G — plus the C→G transversion with a CGBE. Any other change, and every indel, needs prime editing; the tool says so rather than offering a guide that cannot work.

  • The off-target search is not genome-wide. It sees only the sequences you paste in. Pair it with a genome index before committing to a guide.
  • The activity score is a heuristic assembled from published sequence rules, printed term by term. It is not Rule Set 2 / Azimuth and predicts no indel frequency.
  • Cut coordinates are the canonical geometry. Real ends are heterogeneous, especially for Cas12a.
  • "Silent" is only checked when a reading frame is supplied, and only against the codon table — not against splice sites, regulatory elements or codon usage.
  • Base-editing efficiencies are positional heuristics, not a trained model, and window edges are soft.
  • Chromatin state, delivery method and cell type move real editing efficiency more than any sequence rule. Always validate two or three guides experimentally.

CRISPR diagnostics (SHERLOCK / DETECTR)

Applies to: CRISPR Diagnostics — SHERLOCK / DETECTR

Cas12a and Cas13a do something no hybridisation probe does: once they recognise their target they begin cutting any single-stranded nucleic acid in the tube. Supply a quenched reporter and that collateral activity becomes the readout. Method: Gootenberg et al. (Science 2017) and 2018 for SHERLOCK, Chen et al. (Science 2018) for DETECTR.

  1. Pre-amplification. Collateral detection alone reads picomolar; with isothermal pre-amplification it reads femtomolar to attomolar. RPA runs at 37–42 °C, so for Cas12a and Cas13a it needs no thermal cycler at all.
  2. The crRNA. A Cas12a spacer behind a TTTV PAM inside the amplicon, or a 28-nt Cas13a spacer against the transcript. Where a variant is marked, one allele-specific crRNA per allele.
  3. The reporter. A quenched ssDNA (Cas12a) or ssRNA (Cas13a) oligo, read by fluorescence or on a lateral-flow strip.

All three come out of a single run, together with a protocol sketch and the controls that make the result interpretable.

  • One target sequence, 5′→3′, FASTA header optional. Only the first record is used.
  • Mark the variant to genotype as [C/T]; here the two alleles are simply the two things to tell apart, with no "reference vs edit" asymmetry.
  • Leave the brackets out for presence/absence detection, and the crRNAs are tiled across the sequence instead.
  • Paste 400–600 nt around the site. RPA primers are 30–35 nt and must stay clear of the crRNA target, so a short paste leaves nowhere to put them.
  • For an RNA target (a virus, a transcript) paste the cDNA sequence and use RT-RPA at the bench — the design is the same.

  • DETECTR — LbCas12a or AsCas12a (dsDNA, TTTV PAM, 37 °C). Targets the amplicon directly, no transcription step. Needs a TTTV PAM in the right place, which is the main constraint on where the assay can sit.
  • SHERLOCK — LwaCas13a (RNA, no PAM, 37 °C). No PAM means the crRNA can be placed almost anywhere, which is what makes single-base discrimination practical — but the amplicon must be transcribed, so one RPA primer carries a T7 promoter.
  • One-pot — AapCas12b (dsDNA, TTN PAM, 60 °C). Thermostable, so it can share a tube with LAMP for a genuine single-step assay. Its sgRNA scaffold varies between constructs, so the tool gives the spacer and tells you to append your supplier's scaffold rather than inventing one.

Design the LAMP primer set for a one-pot Cas12b assay in the LAMP tool against the same region, and place the crRNA target between F2 and B2 so it survives in the dumbbell structure.

Sites are scored on GC, homopolymers, linguistic complexity, self-structure, pairing against the direct repeat, and a local accessibility proxy; for Cas13a the protospacer flanking site (PFS) is reported too, and a G there is penalised. The output gives the crRNA as RNA and, separately, the DNA template with a T7 promoter for in-vitro transcription.

Telling two alleles apart is not a matter of the variant merely being covered — it has to sit where the enzyme reads it:

  • Cas13a: spacer position 3, with a synthetic mismatch planted two bases away. The deliberate second mismatch makes even the on-allele duplex marginal, so the off-allele duplex — carrying two mismatches — fails outright. This is the design from Gootenberg et al., and it is what turns a detector into a genotyper.
  • Cas12a: inside the PAM-proximal seed, where a single mismatch already blocks cleavage.

One crRNA is produced per allele. Run them in separate wells and call the genotype from which one fires — never from a single well, because a failed reaction and a negative allele look identical. If no site can place the variant where the platform discriminates, the tool falls back to sites that merely cover it and says clearly that allele discrimination will be weak or absent.

  • RPA primers are not selected on Tm. Recombinase loading, not duplex stability, is what limits the reaction, so the rules are compositional: 30–35 nt, balanced GC, no 5′-terminal G, a G/C anchor in the last three bases, no homopolymer longer than five. Tm is printed for information only.
  • Amplicons are short — 100–200 bp amplifies fastest — and must contain the crRNA target with a margin either side, so that no primer competes with the crRNA for the same bases.
  • The pair is checked for inter-primer complementarity, which in RPA is a more common failure than mispriming.
  • For Cas13a the tool says which primer to tag. The T7-tagged primer determines which strand becomes RNA, and the crRNA only works against one of them; get this backwards and the assay is silent however good the crRNA is.

RT-RPA handles RNA input in the same tube — add a reverse transcriptase, no separate cDNA step.

  • Cas12a chews single-stranded DNA and prefers poly-T; Cas13a chews RNA and, in the Lwa orthologue, prefers uridine. The tool gives the matching quenched reporter in ordering syntax.
  • Fluorescence (FAM / Iowa Black FQ) on any plate reader or qPCR instrument; lateral flow (FAM / biotin) on a dipstick, read at 2–5 min before the strip over-develops.
  • Multiplexing needs different enzymes, not different crRNAs: collateral cleavage is indiscriminate, so two crRNAs with the same enzyme in one tube cannot be told apart. SHERLOCKv2 multiplexes by pairing orthologues with different collateral preferences and spectrally separated dyes.
  • Controls, every run: a no-template control through the full workflow including amplification; a no-crRNA control, which separates genuine collateral activity from reporter degradation; and a positive control at a known copy number to anchor the time-to-signal. For genotyping, both alleles against both crRNAs — the answer is the ratio, not one well firing.

Time-to-threshold varies roughly with the log of input copies, so a standard curve makes the assay semi-quantitative; end-point fluorescence saturates and does not.

  • Nothing here predicts a limit of detection. LoD depends on the pre-amplification, the enzyme lot, the reporter concentration and the sample matrix; it has to be measured.
  • Single-base discrimination is a designed property, not a guaranteed one. How marginal the on-allele duplex ends up depends on the flanking sequence. Titrate both alleles.
  • Accessibility is a local hairpin proxy over 60 nt, not a folding calculation — fold the amplicon in the Structure Folding tool before committing.
  • Protocol numbers are starting points, not validated conditions.
  • Carry-over contamination is the practical failure mode of every amplification-coupled CRISPR assay. Separate the pre- and post-amplification benches physically.

Gibson Assembly primer design

Applies to: Gibson Assembly Primer Design

Gibson assembly enables seamless joining of multiple DNA fragments in a single isothermal reaction using a 5′ exonuclease, DNA polymerase, and DNA ligase. Fragments must share ≥20 bp homology with adjacent segments (overlap Tm ≥50°C).

  1. Arrange all DNA fragments in the intended final assembly order in a single FASTA file. Include vector sequence as the first and last entries for circular construct design.
  2. Optionally, paste any pre-existing primers in the My Primers tab to avoid duplicating them in the new set.
  3. Adjust primer length, Tm range, and 3′-end pattern, then click Generate.
  4. Review the Report (Primer list) tab — including its # Notes — for individual primers, and the PCR primer pairs tab for every amplification reaction plus the vector-anchored pair that amplifies the assembled construct.

  • Sequences must be in FASTA format, listed in the desired assembly order.
  • Adjacent fragments must share an overlap of ≥20 bp with a Tm ≥50°C to ensure efficient exonuclease digestion and annealing.
  • For circular constructs, include vector (backbone) sequence at both the first and last positions. The tool will automatically generate primers spanning the junction.
  • Fragment names (FASTA headers) are used in the output to label each primer pair — use meaningful names.

  • Min. length (nt): minimum length of the annealing part of the primer (default 15; the overlap tail is added on top of it). There is no maximum length — the primer is extended as far as needed to reach the minimum Tm. If reaching the minimum Tm would push the Tm above the maximum, the primer is instead shortened (even below the minimum length) so that it stays under the Tm ceiling; these cases are flagged in the report Notes.
  • Tm range (°C): the maximum Tm is an enforced ceiling — no primer ever exceeds it. The minimum Tm is the target, and the shortest primer reaching it is selected. Default 60–62°C.
  • Min. linguistic complexity (%): rejects low-complexity / repetitive primers (≥70% recommended). If no primer can be found at the requested value, the tool automatically relaxes it step-by-step (down to 10%) and notes the value actually used.
  • 3′-end pattern: use N for any nucleotide, or specify patterns like WSS to enforce a particular 3′-terminus composition.

  • Report (Primer list): all designed primers with name, sequence, length, Tm, GC%, and LC%. A # Notes block at the end flags every fragment where the minimum complexity was relaxed, where the primer was shortened below the minimum length to respect the Tm ceiling, or where no primer could be found — together with the likely reason.
  • PCR primer pairs: the forward and reverse primer for each amplification reaction. The annealing portion is in lowercase and the 5′ overlap tail (reverse-complement of the neighbouring fragment end) is in UPPERCASE.
  • Assembled-construct amplification pair: a separate, clearly labelled block at the bottom of the PCR primer pairs tab — see the next panel.

At the bottom of the PCR primer pairs tab the tool reports one extra vector-anchored primer pair for amplifying or verifying the whole joined cassette after assembly:

  • A forward primer placed anywhere on the first (left) vector arm and a reverse primer anywhere on the last (right) vector arm, both pointing inward across the inserts.
  • The position is free — it slides along the arm to find a good primer — so the primer is not forced onto the palindromic cloning site at the fragment end. Where only a short primer is dimer-free, it is reported with a note (it may be below the minimum length).
  • An approximate amplicon size is given (left-arm reach + inserts + right-arm reach), with a warning if the two primers might form a primer-dimer.

Tip: provide a sufficiently long vector sequence as the first/last entries. If a vector arm is only a short polylinker of palindromic restriction sites, dimer-free primers may be short or unavailable — extend the vector sequence or move the fragment boundary.


Multiplex tiling PCR panel design

Applies to: Custom multiplex tiling PCR panel design tool

  • Splits a target region into overlapping amplicons (tiling) and designs primers for sequencing panels.
  • Generates two complementary pools (Panel A and Panel B) to reduce primer competition across adjacent amplicons.
  • Reports primer lists and ready-to-use panel mixes for multiplex PCR workflows.

  • Target sequence(s) in FASTA; mark the region with […] and exclude problem areas with /.../ (see Targeting markup).
  • Amplicon size range: choose by platform (e.g., shorter for Illumina; longer for ONT).
  • Gap between amplicons: set to 0 for full coverage; increase for fewer primers.
  • Optional pre-designed primers: provide a list for compatibility with existing assays.
  • Non-specific priming control (default on) — keeps primers off dispersed/interspersed repeats elsewhere in the flanking sequence.
  • Low complexity priming control (default on) — keeps primers off other SSR/microsatellite and telomere-like stretches in the flanks, the same blocks detected by Repeat Finder. Disable either control only if it prevents a primer from being found and the flanks cannot be replaced.

Each target's block in the PCR Primer Pairs tab opens with a Target coverage line, reporting what fraction of the marked target region (not the flanking primer-search window) is spanned by amplicons:

  • Panel A / Panel B — each panel's own coverage, on its own. Neither panel is designed to reach 100% alone — that is what the other panel's amplicons are for.
  • Combined A+B — the union of both panels' amplicons; this is the number that matters for "did we cover the whole target". 100% combined coverage with two genuinely gapped individual panels is the ideal outcome.
  • If combined coverage falls short, widen the amplicon size range, lower Gap between amplicons, or relax the masking/complexity controls above — the usual reasons a stretch has no valid primer on either panel.

Gator BLI probe design

Applies to: Oligo Probe Design for the Gator® GeneSwift Assay Kit

Designs paired oligo probes for the Gator Bio GeneSwift Assay Kit, which determines AAV vector titer by biolayer interferometry (BLI). A two-step procedure combines lysis and genome hybridization in one tube, followed by BLI detection.

Each assay requires a matched pair of probes that hybridize to the same strand of the target sequence:

Fluorescein-labeled probe
  • 35–40 nt long
  • Labeled only at the 5′ end with fluorescein
Biotin-labeled probe
  • 35–40 nt long
  • Biotins added at both 5′ and 3′ ends
  • Also carries a 30-T spacer (poly-T tail)
The two probes must not overlap. Use the probe distance parameter to control the gap between them.

  • Pick any region within the insert except the Inverted Terminal Repeats (ITRs).
  • For genome integrity assessment, target the ends of the insert, e.g., the CMV enhancer at one end and SV-40 at the other.
  • Use [] directly inside the pasted sequence to pin each probe's location individually.
  • Use // to exclude problematic regions (repetitive elements, secondary structure hotspots) from probe placement; can be repeated multiple times.

  • Probe length (35–40 nt): length of each individual oligo probe.
  • Probes distance (>5 nt): minimum gap (in nucleotides) separating the two probes along the target strand. The tool also reports pairs at the configured separation.
  • Dimer screening: no free-energy value is computed or reported. Each candidate is rejected if it self-dimerises, and a pair is reported only when the two probes do not cross-dimerise with each other.
  • Linguistic Complexity (LC%): measures sequence vocabulary richness; 100% = maximum diversity. Higher values reduce non-specific hybridization risk.

  • Paste or upload sequences in FASTA format; name is optional but recommended (e.g., >EGFP).
  • Retrieve a sequence directly from NCBI by entering a nucleotide accession ID (e.g., A02710) and clicking Retrieve Sequence.
  • The name and sequence string can be separated by either a space or a tab; style must be consistent for all entries.
  • Input is case-insensitive; only standard IUPAC nucleic acid characters are accepted.

  • Results are reported in two tabs, Forward probes (positive strand) and Reverse probes (negative strand).
  • Pairs are listed in order of position along the target; test several of them to find the best performer experimentally.
  • Submit the final probe sequences to an oligo manufacturer (e.g., Eurogentec, LGC Biosearch Technologies, Eurofins) with the appropriate fluorescein and biotin modifications specified.

PrimerAnalyser

Applies to: PrimerAnalyser — comprehensive single-oligo analysis

Provides comprehensive analysis of a single oligonucleotide sequence written with standard bases, degenerate IUPAC codes, uracil (U), inosine (I) and locked nucleic acid residues (E/F/J/L). Results update instantly as you type.

  • Melting temperature Tm (°C) — reported twice. The first line is the nearest-neighbour Tm, the one to use for reaction design, and it names the parameter set it used. The line below it is the classical empirical Tm = 77.1 + 11.7·log[K⁺] + (41·(G+C) − 528)/L, which ignores primer concentration and folds Mg²⁺ into a single crude salt term; it is there for comparison with older calculators and routinely differs by several degrees.
  • GC content (%)
  • Linguistic complexity (LC% and YR%)
  • Length (nt) and base composition — every residue is counted as the base it occupies, so the four columns add up to the length — except for the three-fold degenerate codes B, D, H and V, which add 0.333 to each of three columns, so a sequence containing them totals slightly under its length. Locked (LNA) bases count as their parent base (E to A, F to C, J to G, L to T) and contribute to GC% accordingly; uracil is counted under T and inosine under G. Degenerate IUPAC codes are counted fractionally (for example R adds 0.5 to A and 0.5 to G). Molecular weight still uses each residue's own mass, so U, I and LNA bases are weighed correctly even though they share a composition column.
  • Melting temperature — which parameter set is used. A sequence written with U and no T is treated as RNA and scored with the RNA/RNA nearest-neighbour parameters of Xia et al. 1998, including their single per-duplex initiation term and the penalty for each helix end closed by an A–U pair. Anything else is scored as DNA with the Allawi & SantaLucia 1997 parameters; locked (LNA) residues there add the McTigue et al. 2004 increments. A sequence mixing T and U is ambiguous and stays on the DNA path.
    Three limits worth knowing: the salt correction is Owczarzy's, which was parameterised on DNA, so RNA melting temperatures away from ~1 M Na⁺ are an approximation; and a step in which both residues are locked has no published parameter at all, because McTigue measured one locked position at a time. Those steps are filled by applying both of his increments additively — his model carried one step further rather than measured values — so melting temperatures for runs of consecutive LNA residues are approximate. Third, a locked or inosine residue in a sequence that has taken the RNA path contributes nothing at all: the Xia table covers only A, C, G and U, so every nearest-neighbour step touching an E/F/J/L or an I scores zero. An oligonucleotide written with U and LNA letters is therefore under-scored, and the citation line names Xia alone with no hint that the locked residues were ignored.
  • Self-dimers, with Tm and ΔG for each duplex
  • Extinction coefficient ε (L/mol·cm) — a nearest-neighbour estimate from tables that cover A, C, G and T only, so uracil is scored as thymine and locked (LNA) bases as their parent base; inosine is scored as guanine, whose ε is higher, making that one an upper bound rather than a measured value. A sequence shorter than 4 nt returns no ε at all — it is reported as 0, and the OD260, µg and nmol figures below it are then not computed.
  • Molecular weight (g/mol)
  • Amount per OD unit (nmol/OD260)
  • Mass per OD unit (µg/OD260)
  • Dilution and resuspension calculator

  • Paste a single oligo sequence (5′→3′) with or without a FASTA name header.
  • Standard IUB/IUPAC degenerate codes are supported — see Input formats → Allowed nucleotide codes.
  • U = Uracil (RNA), I = Inosine.
  • LNA codes: E=LNA-dA, F=LNA-dC, J=LNA-dG, L=LNA-dT.
  • Tm calculations use nearest-neighbor thermodynamic parameters, accounting for primer concentration, salt, and Mg²⁺.

Three calculators are available:

  • Dilution: calculate stock volume needed to reach a target working concentration and volume.
  • Resuspension: calculate how much TE/water to add to a lyophilised oligo to reach a desired stock concentration.
  • Amount from OD/mass/nmol: inter-convert between OD260, mass (µg), and nanomoles using the oligo's extinction coefficient and molecular weight.

  • Primer concentration (0.01–5 µM): concentration of the oligo used in your reaction.
  • Total salt concentration (Na⁺, K⁺, NH₄⁺, Tris⁺; 1–1000 mM): matches your buffer conditions.
  • Mg²⁺ concentration (0–10 mM): free magnesium concentration in your reaction.
  • Dimer sensitivity (1–10): controls the detection threshold for self-dimer reporting.

PrimersList

Applies to: PrimersList — batch primer analysis

Analyzes multiple primers simultaneously in a single run: Tm, GC%, secondary structures (hairpins, self-dimers, cross-dimers), linguistic complexity, molecular weight, extinction coefficient, and OD calculations.

Accepts two equivalent formats:

FASTA format
>m13-47
cgccagggttttcccagtcacgac
>RP
tttcacacaggaaacagctatgac
Name + sequence (Excel-friendly)
m13-47  cgccagggttttcccagtcacgac
RP      tttcacacaggaaacagctatgac

The Tm Calculation Parameters card sets the conditions every melting temperature on this page is computed under:

  • Primer concentration (0.01–5 µM, default 0.20).
  • Total salt — Na⁺, K⁺, NH₄⁺ and Tris⁺ combined (1–1000 mM, default 55).
  • Mg²⁺ concentration (0–10 mM, default 1).
  • Dimer sensitivity (1–10, default 4) — despite the name this is a stringency: the shortest base-paired stretch reported as a dimer is the value plus one, clamped to 4–9 nt. Raise it to see only the longer duplexes, lower it to catch marginal ones.

Leave the first three at their defaults (0.2 µM / 55 mM / 1 mM) to keep Tm comparable with the design tools, which fix them — see Common parameters → Tm calculation conditions.

  • Features: tabular summary — Tm, GC%, LC%, length, base composition, Mw, extinction coefficient, OD values — one row per primer.
  • Dimers: self-dimer and cross-dimer analysis for all primer combinations, ranked by binding strength.

PrimersList handles specialized oligonucleotide types:

  • Molecular beacons — stem-loop probes; the stem sequence lowers the apparent Tm and affects dimer analysis.
  • Scorpion primers — self-probing amplification primers with an internal probe region.
  • Standard degenerate primers, RNA oligonucleotides (U), and LNA-modified sequences (E/F/J/L).

DNA Sequence Utilities

Applies to: DNA Sequence Utilities

Two families of operation on one page. Operations are small, independent DNA text transforms — reverse complement, bisulfite conversion, FASTA ⇄ column reformatting, clustering by homology — applied to every record independently. File formats, annotations & reads converts between sequence file formats, pulls features and coding sequences out of annotated records, and runs quality control on FASTQ reads. Paste or upload your data, pick an operation, and the result is ready to copy or download.

The two families treat the input differently, and that difference matters. Every Operation first cleans each record to valid IUPAC nucleotide letters, which discards headers, gaps, digits and quality strings — exactly what you want before reverse-complementing, and exactly what a converter must not do. The File formats card therefore parses the file as written and preserves case, gap characters, FASTQ quality strings and the whole feature table.

The Operations card accepts any of the formats the site's other tools use — no need to reformat first:

  • FASTA (>name header lines).
  • Name + sequence, one record per line, separated by a tab or a space (Excel-friendly).
  • A single bare sequence with no header at all.

The File formats, annotations & reads card reads all of those and, in addition:

  • FASTQ — reads with quality strings, Phred+33 or Phred+64. The usual four-line layout and the rarer wrapped layout are both accepted.
  • GenBank flat file (LOCUS//), including the FEATURES table.
  • EMBL flat file (ID//), including the FT feature table.
  • Clustal (.aln) — the conservation line is recognised and skipped.
  • PHYLIP — interleaved or sequential, strict 10-character or relaxed names; the layout is worked out from the declared ntax nchar dimensions.
  • NEXUS — the DATA or CHARACTERS block, interleaved or not; [comments] and a TAXA block are ignored.
  • CSVname,sequence per line.

Leave Read as on Auto-detect the input format unless the guess is wrong — the format is identified from the first meaningful line. Setting it explicitly is useful for a file with an unusual header, or to force a strict reading.

  • Reverse Complement — the sequence of the opposite strand, read 5′→3′ (what you'd order as a reverse primer).
  • Complement (no reverse) — complements each base in place without reversing the order.
  • Reverse (no complement) — reverses the letter order only; rarely needed on its own, included for completeness alongside the two above.
  • UPPERCASE / lowercase — case conversion only, letters are otherwise untouched.
  • C→T Bisulfite (both strands)in silico bisulfite conversion of both strands, reported as an aligned duplex rather than plain FASTA: the converted sense strand 5′→3′ above the converted antisense strand 3′→5′. Same rule as the PCR/qPCR/KASP tools: only cytosines not followed by guanine (non-CpG context) become thymine; CpG sites are preserved. The two strands are no longer complementary once converted — that asymmetry is the point of the experiment, not a display error.
  • FASTA → Column — converts to one name<TAB>sequence row per record (paste straight into Excel).
  • → FASTA — converts column or bare input back into standard >name / sequence FASTA records.
  • Sort by Length — reorders multi-record input, longest sequence first.
  • Length & CG% report — a quick per-record table, no sequence transform.
  • Cluster Sequences — groups the input records into homology clusters; needs at least 2 sequences. See Cluster Sequences by Homology below.

All input is cleaned to valid IUPAC nucleotide letters first (see Input formats → Allowed nucleotide codes); a record that has no valid bases after cleaning is dropped.

Choose what to write as and press Convert. Output is available in FASTA, FASTQ, GenBank, EMBL, Clustal, PHYLIP (interleaved or sequential), NEXUS, tab-delimited, CSV and GFF3. The Download button names the file after the format it contains (converted.gb, converted.aln, and so on).

  • Line width — characters of sequence per line; also the block width in Clustal, PHYLIP and NEXUS.
  • Letter case — leave as written, or force upper or lower case.
  • PHYLIP namesrelaxed keeps the full identifier (what RAxML, IQ-TREE and PhyML expect); strict truncates to the 10 characters of the original specification and de-duplicates any names that collide as a result.
  • Keep the description — write the whole header line, or just the first token (the identifier). Turning it off is the quickest way to get clean names for a downstream program that splits on whitespace.
  • Pad unequal lengths with gaps — Clustal, PHYLIP and NEXUS are alignment formats and require every row to be the same length. With this on, shorter records are padded with - and a note says how many were padded; with it off, an input whose records differ in length is refused rather than silently altered.

What survives a conversion, and what does not

  • Sequence, identifiers, descriptions, gap characters and letter case are always preserved.
  • GenBank ⇄ EMBL keeps the feature table, locations and qualifiers, plus accession, organism and taxonomy.
  • Converting an annotated record to FASTA, PHYLIP or any other plain format necessarily drops the annotation — that information has nowhere to go. Extract what you need first (below), or write GFF3 alongside the FASTA.
  • Writing GenBank or EMBL from plain input invents the minimum a valid record needs: a source feature spanning the sequence and placeholder division and version fields. It is a well-formed record, not an annotated one.
  • Writing FASTQ from input that has no quality scores fills them with a constant value and says so. Those numbers are placeholders, not measurements — never report them as quality.
  • GenBank records that store their sequence as a CONTIG join instead of an ORIGIN block contain no bases to convert; the note explains this rather than producing an empty record silently.

Pulls annotated regions out of a GenBank or EMBL record — genes, coding sequences, exons, UTRs, repeats and everything else the submitter marked up. Press Load GenBank Example for a worked record, or fetch a real one with View GenBank → Result above and paste it into the Input tab.

  • Feature type — extract everything the record declares (all features, the default) or just one class: CDS, gene, mRNA, exon, intron, tRNA, rRNA, ncRNA, 5′UTR, 3′UTR, regulatory, repeat_region, misc_feature or source. It governs all three buttons below; if the record carries no feature of the chosen type, the note lists the types it does contain.
  • Extract → FASTA — one record per feature. The header carries the source record, feature type, gene or product name, the original location string and the length.
  • Feature table (TSV) — one row per feature with the source record, type, start, end, strand, segment count, length, /gene, /locus_tag, /product, /protein_id and the original location string. Paste straight into Excel.
  • Feature list (GFF3) — the same features as a GFF3 file for a genome browser; a spliced feature becomes one line per segment sharing a Parent.

Locations

The full INSDC location grammar is understood: a single base (467), a span (340..565), the opposite strand (complement(…)), a spliced feature (join(…) and order(…)), partial ends (<1..888, 1..>888), uncertain and between-base positions (102.110, 123^124) and any nesting of these. Segments are assembled in biological order, so complement(join(a,b)) gives the reverse complement of the joined product rather than the pieces in file order. A location pointing at another accession (J00194.1:100..202) cannot be sliced from the record in front of you; those features are skipped and counted in the note.

Translation

  • Tick Translate CDS to protein to translate every extracted CDS. The reading frame comes from the feature's own /codon_start, and the code from its /transl_table; the Genetic code selector is the fallback for a CDS that declares neither. All 25 NCBI translation tables are available.
  • A codon containing an ambiguity code is translated when every base combination it stands for gives the same residue (so CTN is L) and to X otherwise. A codon containing a gap is X — gaps consume a codon position rather than being deleted, so the frame downstream is never silently shifted.
  • The result is compared against the record's own /translation qualifier where one exists. A mismatch is flagged in the header as differs_from_/translation, which usually means a /transl_except (selenocysteine, a non-standard initiator) or an in-record correction that a plain translation cannot reproduce. An internal stop codon is flagged as internal_stop; in an ordinary nuclear gene that is a sign the frame or the genetic code is wrong.
  • A trailing stop codon is not written into the protein, matching the convention used by GenBank's own /translation.

Quality report summarises a read file; FASTQ → FASTA drops the quality strings, optionally filtering and trimming on the way. Press Load FASTQ Example to see both on a handful of reads.

The report covers

  • Read and base counts, read-length range and mean, GC content and the fraction of ambiguous bases.
  • The quality encoding, detected from the lowest score present — Phred+33 (Sanger and Illumina 1.8+) or Phred+64 (Illumina 1.3–1.7). The decision is made from the lowest quality character present: below ‘;’ the file can only be Phred+33, at ‘@’ or above it is read as Phred+64. A read set filtered so that no base falls below Q31 carries no character low enough to identify it and is reported as Phred+64 — if that label looks wrong for a modern Illumina file, every Q value below it is 31 too low.
  • Mean quality, the percentage of bases at or above Q20 and Q30, and the percentage of reads whose mean is at or above Q30.
  • Mean quality per position and base composition per position, as text bar charts. Long reads are binned so the table stays readable; the bin size is stated in the heading.
  • The read-length distribution, and the fraction of reads that are exact duplicates of another read.
  • An adapter screen: an exact search for the Illumina TruSeq/universal, Nextera and small-RNA adapters, and for poly-G (the two-colour no-signal artefact) and poly-A.

Filtering and trimming

These three boxes apply to FASTQ → FASTA only. The Quality report always describes the file exactly as it stands, so run the report first, then filter.

  • Min length — discard reads shorter than this, applied after trimming.
  • Min mean Q — discard reads whose mean Phred score is below this.
  • 3′ trim < Q — walk in from the 3′ end while the score is below the threshold and cut there, the classic quality-trim used before assembly. Leave at 0 for no trimming.

Because everything runs in the browser, a whole sequencing run is not a realistic input — the report samples the first 200 000 reads and says so when it does. That is ample for judging run quality; use it to decide what to do, not as a substitute for a pipeline-level report. Nothing is uploaded: the reads never leave the page.

Fills the Input tab with a fresh random sequence (uniform A/C/G/T composition, no biological meaning) of the given length — default 1000 nt, up to 1,000,000 nt. Useful as quick scratch/test input, e.g. to sanity-check another operation or a downstream tool. This replaces the current Input tab contents, like Load Example.

Groups the input records into homology clusters by the fraction of shared 12-mers — the overlap coefficient of each pair's distinct 12-mer sets, taken relative to the shorter sequence so a length difference doesn't deflate it. Shared long k-mers stay high for genuinely homologous sequences even across substantial divergence and indels, and are ~0 for unrelated sequences, so this reliably separates homologs from noise. Both strands are compared. Clustering is single-linkage (A~B and B~C group A, B and C together). Needs at least 2 sequences. Each cluster is reported as a labelled FASTA block, one record per member, ready to copy straight into another tool or an alignment.

  • Shared k-mers % (k = 12; default 30, range 5–95) — the minimum percentage of shared 12-mers (relative to the shorter sequence) required to place two sequences in the same cluster. Homologous genes typically share 30–60%; unrelated sequences ~0%. Lower it to catch more diverged homologs, raise it to be stricter. Note this is a shared-k-mer fraction, not a literal alignment percent-identity — a pair at ~85% identity with indels may share only ~20–30% of its 12-mers.
  • Sequences that don't match anything closely enough are listed separately under Unclustered, not forced into a cluster.
  • A member flagged reverse-strand match (noted on its header line) matched the rest of its cluster on the opposite strand — reverse-complement it before aligning by eye against its cluster-mates; the sequence itself is still reported exactly as given, not auto-flipped.
  • Normalising by the shorter sequence means a length difference on its own does not weaken the match (a short homolog still clusters with a much longer one). It compares whole records as given without aligning or trimming, so extra unrelated flanking sequence on the shorter member — which adds non-shared 12-mers — can still lower the score.

Enter a NCBI nucleotide accession (e.g. NC_045512, NM_000546) and choose what to retrieve:

  • Fetch FASTA → Input — loads just the sequence into the Input tab, ready for any operation above.
  • View GenBank → Result — loads the full GenBank flat-file record (features, annotations, references) into the Result tab to inspect or copy. Copy it into the Input tab to run feature and CDS extraction or convert it to another format.
  • Download GenBank (.gb) — saves the same record straight to a .gb file, without going through the Result tab.

This is the one part of the page that talks to a server: the accession is sent to NCBI to fetch the record. Everything else — every transform, conversion, extraction and quality report — runs entirely in your browser.


RNA & DNA Secondary Structure

Applies to: RNA & DNA Secondary Structure

Predicts the structure a sequence folds into and the free energy it releases doing so — the hairpin that sequesters a primer's 3′ end, the stem a LAMP loop primer has to invade, the fold of an sgRNA scaffold, the duplex two primers form with each other. Paste one sequence to fold it on itself, or two to let them interact. The prediction is a minimum-free-energy fold: of all the structures the sequence could adopt, the one with the lowest free energy under the chosen conditions.

A folded molecule is decomposed into loops, and the energy of the whole is the sum of the loops: stacked pairs, hairpins, bulges, interior loops, multiloops and the exterior loop. Each contribution comes from a measured nearest-neighbour table.

  • RNA — Turner 2004. Mathews et al. PNAS 2004;101:7287, and the Turner nearest-neighbour database (Turner & Mathews NAR 2010;38:D280).
  • DNA — Mathews / SantaLucia 2004. SantaLucia & Hicks Annu Rev Biophys Biomol Struct 2004;33:415, with the DNA loop parameters of Mathews et al.

These are the complete published tables, not a reduced model. 1×1, 2×1 and 2×2 interior loops carry their own measured matrices rather than a generic size penalty; hairpins of three, four and six bases are looked up in the tri-, tetra- and hexaloop bonus tables, which the published RNA set alone provides — the DNA set carries no measured special-hairpin bonuses, so a DNA hairpin always takes the generic size term plus its terminal mismatch; loops longer than 30 bases are extrapolated with the usual logarithmic term. Terminal mismatches, dangling ends and the terminal-AU penalty are all applied, using the model in which a helix opening into an exterior or multi loop always takes the terminal mismatch of its two flanking bases.

Both parameter sets are stored as free energy at 37 °C together with the corresponding enthalpy, so changing the temperature rescales every term rather than merely relabelling the result. The recursions and loop energies follow the standard implementation, and the output was checked base pair by base pair against ViennaRNA (RNAfold, RNAcofold, RNAduplex and RNAsubopt) over several hundred random and real sequences of both kinds, at temperatures from 5 to 90 °C: the free energies agree exactly.

  • Sequence. Plain sequence or FASTA; header lines and anything that is not a letter are ignored. T and U are read as the same base, so a DNA sequence can be folded with the RNA parameters and the other way round.
  • Nucleic acid. Detect from sequence reads a sequence containing U and no T as RNA and everything else as DNA. Override it when you mean something else — a DNA oligo whose behaviour you want to compare against its RNA counterpart, for instance.
  • Temperature. The whole parameter set is rescaled by enthalpy. This is the setting that makes the tool useful for PCR: fold a primer at its annealing temperature, not at 37 °C, and a hairpin that looks alarming on paper often turns out to be gone by the time the reaction reaches 60 °C.
  • Suboptimal, kcal/mol. Leave at 0 for the single best structure. Set it to 1–3 to also list every structure within that much of the optimum. A sequence whose second-best structure is 0.2 kcal/mol behind the first has no single well-defined fold; one whose nearest competitor is 4 kcal/mol behind does.
  • Max structures. How many suboptimal structures to keep — 50 by default, up to 1000. The enumeration stops there even if more structures lie inside the window, and it bites at the shipping settings: both the tRNA and the primer-hairpin example fill the list before the window is exhausted.
  • Watson–Crick pairs only. Switches off G·U pairs. In RNA the wobble pair is real and should normally be left on. In DNA the corresponding G·T pair is far weaker than the parameter set implies, so turning it off is often the more realistic choice for oligonucleotide work.
  • Running it. Nothing is computed until you press Predict structure; the example buttons only load the sequences. A result already on screen is not cleared when you edit a sequence, so run it again after any change — otherwise the structure you are reading may belong to the previous input.
  • The worked examples. Four buttons load a sequence and stop there, so the fold is still yours to run. tRNA is a folded RNA with the cloverleaf every textbook draws. Primer hairpin is a polylinker that swallows its own 3′ end. Primer dimer is two different primers meeting. Primer self-dimer is one primer against a second copy of itself — the everyday question, and the interesting case, because this primer also folds into a hairpin, so the two compete: at 37 °C the answer is a hairpin on each strand joined by four pairs, ΔG −9.90 kcal/mol. It asks for together explicitly, so it lands the same way whatever the mode was set to; switch it to hybridisation only and the same pair of strands gives a plain 14-pair duplex at −7.20 instead.

Folding costs time proportional to the cube of the length, so the page stops at 2000 bases. A few hundred bases finish in about a second. The work runs in a background thread, so the page stays responsive while it computes — unless the browser refuses to start one, which opening the page straight from a local file does in some browsers. The same calculation then runs on the page's own thread, and the page sits still until it finishes.

The two modes answer different questions, and the difference between their energies is itself informative — it is the price the molecules pay to unfold themselves before they can pair.

  • Together lets each strand fold on itself as well as onto the other, so a hairpin competes with the duplex. This is the honest picture of what happens in the tube.
  • Hybridisation only forbids pairs within a strand and reports the best duplex the two can form. Use it to ask how good the intended pairing could possibly be, ignoring the competition.

A duplex initiation term is charged once, and only when the two strands actually pair with each other — two sequences that ignore one another come out at exactly the sum of their separate folding energies, so a result of 0 means no favourable interaction rather than a failure to look.

The Linear view tab lays the two strands out side by side, the way a dimer is written on paper. It appears whenever the strands touch at all — including the usual case where each also folds on itself, which is where the radial drawing is hardest to read.

  • Drawing. The classical stems-and-loops picture: helices are straight ladders, loops are circles. Watson–Crick pairs are drawn in teal and G·U / G·T wobble pairs in amber; in a two-sequence result the second strand is red and the gap between the molecules is opened up. Download it as SVG for a figure.
  • Arc diagram. The sequence on a line with every pair as an arc. Nested arcs are a helix, and the picture stays readable at any length, which the drawing does not. It downloads as SVG too.
  • Structure & energies. The dot-bracket notation under the sequence — matching brackets are a pair, a dot is unpaired, and “&” marks the strand break — followed by a loop-by-loop breakdown. The breakdown is the useful part: it says which loop is costing you and which stack is holding the structure together, and the column sums to the reported free energy.
  • Linear view. Two-sequence results only. The strands lie side by side: a rung is a pair between them, an arc over a strand is a pair inside it — a hairpin, which a two-line ladder cannot draw as a rung — and unpaired bases face each other, so an asymmetric loop shows as a gap on one side. Each strand carries a ruler every tenth base, numbered in its own coordinates, so B1 is B’s 5′ end although B is drawn right to left. The picture is headed with the free energy and the conditions, which makes the downloaded SVG readable on its own; the same layout is repeated underneath as monospace text for pasting into a note. Long duplexes wrap into blocks of 70 columns.
  • Suboptimal. Every structure within the window, best first, up to the Max structures limit. Click one to draw it — the row is marked, and every view follows, including the linear one. In the linear view a structure other than the optimal one is compared against it: each pair this structure has and the optimal does not is drawn violet and thicker, and the caption counts them, which is usually the quickest way to see what the alternative actually does differently. A “+” after the count on the tab means the list was cut short.

What the number does and does not mean

  • A free energy is not a melting temperature. A hairpin at −3 kcal/mol at 37 °C may be entirely absent at 60 °C; refold at the temperature you actually use rather than converting in your head.
  • The prediction is of one structure, the lowest-energy one. Where several structures lie within a fraction of a kcal/mol, the molecule does not choose the first one — it populates all of them. Use the suboptimal list to see whether that is the case.
  • Pseudoknots, G-quadruplexes, and the effects of salt, magnesium and strand concentration are outside this model. For salt-corrected duplex melting temperatures use the Tm calculations in the design tools.
  • Ambiguity codes have no measured parameters and are treated as bases that cannot pair, so a sequence full of N will look unstructured.

Multiple Sequence Alignment (MUSCLE-JS)

Applies to: MUSCLE-JS Multiple Sequence Alignment

Browser-based pairwise and multiple sequence alignment using algorithms inspired by MUSCLE. Supports Needleman–Wunsch (global) and Smith–Waterman (local) for pairwise alignment, plus progressive MSA for multiple sequences. It also cleans pasted GenBank/NCBI listings, flags invalid characters, and can auto-correct reverse-strand DNA/RNA before aligning.

  1. Paste or upload sequences in FASTA format. Two sequences trigger pairwise alignment; three or more trigger progressive MSA.
  2. Select the alignment algorithm (global/local) if prompted, or leave the default for standard DNA/RNA sequences.
  3. For DNA/RNA, leave Auto-correct orientation on to flip any reverse-strand reads before aligning; turn it off to align sequences exactly as entered.
  4. Click Generate and review the visual output (coloured alignment) or switch to FASTA/CLUSTAL export format.
  5. Open the Log/summary tab for identity, score, the local-alignment region, which sequences were reverse-complemented, and any input notices.

The tool can automatically detect and correct sequence orientation (forward vs. reverse-complement) before alignment. This is especially useful when working with Sanger sequencing reads or assembly contigs that may be reported in either direction.

  • Auto-correct orientation (Alignment Parameters) is on by default and applies to DNA/RNA only; protein input is never reverse-complemented.
  • Each non-reference sequence is scored against the first sequence in both orientations; if the reverse-complement scores higher, that strand is used.
  • Corrected sequences are flagged in the Log/summary tab and tagged above the visual alignment, so the displayed strand is unambiguous.
  • Turn the switch off to keep every sequence exactly as entered — useful when the strand is already known and you do not want it changed.

  • Visual: colour-coded alignment shown directly in the browser for quick inspection.
  • FASTA: aligned sequences with gap characters (-), suitable for downstream phylogenetic or variant analysis.
  • CLUSTAL: standard interleaved CLUSTAL format compatible with most alignment viewers.
  • The visual view shows a conservation line beneath each block: * marks a column identical in every sequence, . marks a variable column with no gap, and a column containing a gap is left blank.
  • Output block width controls how many columns are shown per line in the visual output (line wrapping).
  • The Log/summary tab reports identity and, for pairwise runs, the alignment score; for local (Smith–Waterman) it also lists the matched region coordinates in each sequence.
  • Use Copy to copy the alignment to the clipboard, or Save view to export it exactly as shown (colours preserved) as a standalone HTML file.

  • FASTA headers start with >; each header begins a new sequence record.
  • Use standard one-letter alphabets (DNA/RNA or protein). Both upper- and lower-case are accepted.
  • Automatic cleaning: position numbers and whitespace from GenBank/NCBI-style listings are stripped, case is normalised, and gap symbols (.) are converted to -, so you can paste blocks straight from a record viewer.
  • Validation notices: if a sequence contains characters that are not valid for its detected type (DNA/RNA or protein), the offending characters and their positions are listed under Notices in the Log/summary tab.
  • Auto mode selects pairwise alignment for 1–2 sequences and multiple sequence alignment for 3+ sequences.
  • The sequence-count badge on the input tab updates as records are detected.

  • Gap open penalizes starting a gap; gap extend penalizes extending an existing gap.
  • Penalties are typically negative (e.g., -10 open, -1 extend). More negative values generally produce fewer/shorter gaps.
  • Global aligns end-to-end (Needleman–Wunsch); Local finds best-matching subregions (Smith–Waterman).

  • The alignment engine uses dense typed-array dynamic-programming matrices with cached substitution-score lookups, so pairwise alignments of longer sequences run substantially faster than a naive implementation.
  • Orientation checks score each strand without building a full traceback matrix, keeping memory low even on long reads.
  • The progressive guide tree is built with the nearest-neighbor-chain UPGMA algorithm (O(n²) instead of O(n³)), so adding more sequences scales much better.
  • The Time field in the Log/summary tab reports how long the last run took.
  • It remains an in-browser tool: very large inputs (many long sequences) are still bounded by browser memory, since the multiple-alignment step holds full matrices for the longest profile.

  • Privacy: alignments run locally in JavaScript; sequences are not sent to a server by this page.
  • Limitations: a compact, educational implementation inspired by MUSCLE; results may differ from full MUSCLE builds and other aligners.
  • If the output looks wrong, try: (i) switching output format, (ii) simplifying input characters, or (iii) adjusting gap penalties.

Phylogenetic Tree Builder

Applies to: Phylogenetic Tree — NJ / UPGMA from a multiple alignment

Turns a set of DNA or protein sequences into a phylogenetic tree. Unaligned input is aligned first with the same MUSCLE-JS engine as the MSA tool; a pairwise distance matrix is then corrected for multiple substitutions and clustered into a tree. The tree can be re-rooted, rotated, collapsed, relabelled and pruned on screen, and exported as vector graphics or as a standard tree file.

  1. Paste FASTA or a Newick string on the Sequences tab, or press one of the DNA / Protein / Newick example buttons — an example only fills the input box and discards any tree a previous run left on screen.
  2. Set the tree method, distance model, gapped-site handling and bootstrap replicates in Inference parameters.
  3. Press Build tree. Nothing is computed until you do, and the Tree tab opens by itself when the run finishes.
  4. Editing the input afterwards leaves the drawn tree standing — re-rooting, renames and collapsed clades survive — but the status line asks you to press Build tree again.

  • Neighbour-joining (Saitou & Nei 1987) — the default. Does not assume a molecular clock, so branch lengths may differ between lineages. The tree is unrooted, but Midpoint root (NJ) in Inference parameters is ticked by default, so an NJ tree is already drawn midpoint-rooted; untick it to see the raw trifurcation, or re-root on a chosen branch in the Tree tab.
  • UPGMA (Sokal & Michener 1958) — ultrametric: every leaf ends at the same distance from the root. Only appropriate when rates really are clock-like.

Distance models (Auto picks by sequence type):

  • DNA: p-distance (uncorrected) · Jukes-Cantor · Kimura 2-parameter (default; separates transitions from transversions).
  • Protein: p-distance · Poisson correction · Kimura 1983 approximation (default).

Gapped sites: pairwise deletion ignores a column only for the pairs where it is a gap (keeps more data); complete deletion removes any column gapped in any sequence (uniform but stricter).

Non-parametric bootstrap (Felsenstein 1985): alignment columns are resampled with replacement, the tree is rebuilt for each replicate with the same method and model, and each internal branch is labelled with the percentage of replicates that recovered it.

  • 100 replicates is a quick default; 500-1000 is usual for publication.
  • Values above ~70% are normally treated as reasonable support; low values mean the data do not distinguish that grouping — not that the tool failed.
  • Set the count to 0 to skip bootstrapping (much faster on large sets).

Click any node (leaf or internal) to select it, then use the buttons:

  • Re-root here — place the root on the branch above the selected node (outgroup rooting). Branch lengths and topology are preserved.
  • Rotate — swap the order of a clade's children; cosmetic only.
  • Collapse — fold a clade into a triangle; click again to expand.
  • Rename — relabel a leaf or annotate an internal node.
  • Delete — prune a taxon or a whole clade; the remaining branch lengths are merged so distances stay correct.

Layouts: phylogram (x-axis is evolutionary distance, with a scale bar), cladogram (topology only) and circular. Support values, branch lengths, row spacing and width are toggled in the toolbar, which also carries two tree-wide buttons: Ladderize, which sorts every clade by size — the tree is drawn ladderized to begin with, so use this to restore that order after rotating or pruning — and Midpoint root.

  • Input: three or more sequences in FASTA — aligned, or unaligned with Align first if needed ticked (on by default) — DNA or protein; or a Newick string, which must begin with ( and end with ;, to load and edit a tree computed elsewhere. Fewer than three sequences, or sequences of unequal length with the alignment step switched off, stop the run with a message in the status line instead of producing a tree.
  • SVG — vector, exactly as drawn; opens in Illustrator or Inkscape for figure preparation.
  • PNG — raster at 2× for slides and quick sharing.
  • Newick / NEXUS — standard tree files (with support values) for MEGA, FigTree, iTOL and similar.

Exports reflect your edits: what you see is what is written out.

  • Distance-based inference only. Maximum likelihood and Bayesian methods are not included — they need run times and numerical machinery that do not fit a browser tab. For those, export the alignment and use IQ-TREE, RAxML or MrBayes.
  • Saturated pairs are flagged. When sequences are too divergent for the chosen correction (for example p ≥ 0.75 under Jukes-Cantor), the distance cannot be computed. Such pairs are placed beyond the largest measurable distance so a tree can still be drawn, and the run is reported as containing saturated pairs — treat those branch lengths as lower bounds, or switch to p-distance or a protein alignment.
  • Alignment quality dominates. A tree can only be as good as the alignment it was built from; check the Alignment tab before trusting the topology.
  • Orientation is not corrected here. The alignment step runs with orientation auto-correction switched off, unlike in the MSA tool: a reverse-strand sequence is aligned exactly as entered and shows up as a long spurious branch. Orient the input first — for example in the MSA tool with Auto-correct orientation on — and paste the corrected sequences back.

Sequence Assembly (CAP-JS)

Applies to: CAP-JS — sequence and genome assembly

Joins short DNA fragments into one or more consensus contigs by overlap–layout–consensus, the approach of CAP3 (Huang & Madan, Genome Research 9:868, 1999). It is built for the everyday case: a handful to a few hundred Sanger reads, cloned inserts, or contigs from elsewhere, in unknown orientation, with imperfect base calls, and possibly containing repeated sequence. Where the reads genuinely disagree, the disagreement is reported rather than averaged away: a position with two well-supported alleles is written into the consensus as an IUPAC ambiguity code.

  1. Paste or upload your fragments in FASTA on the Fragments tab. Orientation does not matter — reads on the minus strand are detected and flipped, and the report says which ones were.
  2. If you have PHRED quality values (a .qual file), paste them on the Quality tab. They are optional, but they let the consensus weight each base by its quality and let poor read ends be clipped automatically.
  3. Press Assemble. Progress is shown under the buttons; pressing the button again stops the run.
  4. Read the Assembly report first: how many reads were placed, how many contigs came out, and — importantly — which overlaps were rejected or conflicting. Those are the places where the data is ambiguous.
  5. Take the contigs from Consensus (FASTA), and inspect any disagreement in the Alignment view, where each read is shown under the consensus with dots for agreement.
  6. Follow on by checking a contig against a database, aligning several contigs in MSA, or designing primers on the consensus in Primer Design.

The defaults suit Sanger-quality reads of a few hundred to a few thousand bases. Three parameters do almost all the work:

  • Min. overlap (40 nt) — the shortest alignment accepted as evidence that two fragments belong together. Raise it if unrelated fragments are being joined; lower it only if you know your fragments barely overlap.
  • Min. identity (90 %) — lower it for noisy reads; raise it to keep similar-but-distinct sequences (paralogues, gene family members, related strains) in separate contigs.
  • Max. unmatched overhang (30 nt) — the repeat filter. In a genuine join, one fragment runs off each end of the alignment; two fragments that merely share an interspersed repeat both continue past the alignment on both sides. Raising this admits more joins and more mistakes.
  • Seed word length k (12) — how overlaps are found before they are aligned. A smaller k finds more overlaps in error-rich data at the cost of speed; a larger one is faster and more specific. Fragments shorter than k cannot be seeded at all.
  • Min. fragment length (40 nt) — shorter records are reported as dropped rather than silently used.
  • Trim from both ends / Quality clip — fixed or quality-driven clipping of unreliable read ends before assembly. Quality clipping needs the Quality tab.
  • Search both strands — leave on unless every fragment is already known to be on the same strand.
  • Worker threads — see “Speed and threads” below.

When reads disagree at a column, there are two possible reasons and the tool has to tell them apart. An isolated wrong base is sequencing error and should be out-voted. The same minority base seen in several independent reads is polymorphism — a heterozygous site, two alleles of a gene family, two strains in a mixture — and should be preserved.

A column is called polymorphic only when the minority base clears all of:

  • at least Min. reads carrying it (2 by default),
  • at least Min. minor-allele fraction of the column (25 % by default),
  • at least Min. column coverage (3 reads by default), and
  • a statistical test: given the sequencing error rate measured from your own data (it is reported in the Log tab), the chance of that many reads carrying that base by error alone must be below roughly one expected false call per assembly.

The last test is what makes the calls trustworthy at higher coverage. Without it, two coincident errors among eight reads reach 25 % of the column and pass the plain fraction rule — on a 30 kb assembly that produced 23 “polymorphic” sites in sequence that had none.

The consensus then carries the standard IUPAC code: R = A/G, Y = C/T, S = G/C, W = A/T, K = G/T, M = A/C, and B, D, H, V, N for three or more alleles. Every such column is also listed in the report with its coverage, the count of each base, and the code.

Insertions and deletions have no IUPAC letter. A minority indel is therefore reported in the same table as an indel variant, with the number of reads on each side, and the consensus takes the majority.

Repeated sequence is the hardest part of any assembly. If two different places in your target contain the same sequence, a fragment that ends inside one copy looks exactly like a fragment that ends inside the other, and joining them splices two unrelated regions into one contig that looks perfectly confident and is wrong.

Three independent checks guard against this:

  • Overhang — an alignment that leaves a long unmatched tail on both fragments at both ends is a shared repeat, not a join.
  • Repeat content — the tool measures how many times each word occurs across all your fragments, which tells it directly which stretches are single-copy and which are repeated. An overlap that lies almost entirely inside repeated sequence is refused, because nothing distinguishes it from the same repeat elsewhere.
  • Branching — if a fragment's end overlaps several fragments that do not overlap each other, the continuation is ambiguous and the contig stops there.

The consequence is deliberate: the assembly breaks at unresolvable repeats instead of guessing, and the report tells you where. Look at the “Rejected overlaps” and “Layout conflicts” tables — a cluster of rejections mentioning repeat content is exactly the boundary of a repeat. Longer reads that span the repeat with unique sequence on both sides are what resolves it.

If you are assembling something you know has no repeats and the tool is splitting it too readily, lower Min. overlap or raise Max. unmatched overhang, and check the rejection reasons to see which filter is firing.

  • Assembly report — a run summary (reads in, reads placed, singletons, contigs, N50, mean coverage, elapsed time and threads), the parameters actually used, then tab-separated tables: one row per contig, one per read placement, one per polymorphic site, one per overlap used, one per rejected overlap with its reason, and one per layout conflict. Downloadable as text; the tables paste straight into a spreadsheet.
  • Consensus (FASTA) — the contigs, upper-case, with length, reads, cov and variant counts in each header. Unassembled fragments (“singletons”) are appended and marked, unless you turn that off.
  • Alignment view — every read shown under its contig's consensus. A dot means the read agrees with the consensus, so disagreements stand out; - is a gap, +/- after the read name is the strand, and polymorphic columns are marked with * and shown in colour. Downloadable as plain text or as a standalone colour HTML page.
  • Read placements (TSV) — contig, read, strand, start, end and clipping for every fragment, for downstream scripting.
  • Log / summary — the short version, plus the seeding statistics, the measured error rate, and any notices.

Almost all the time goes into aligning candidate fragment pairs, and those alignments are completely independent of one another, so that stage — and only that stage — is spread across Web Workers. Seeding, layout and consensus stay on the page but run in short slices, so the tab keeps responding either way.

Measured on 300 reads of 600–900 nt covering 30 kb: 3.6 s on one thread, 2.4 s on two, 1.8 s on four, 1.5 s on eight. The speed-up flattens because the stages that stay serial take about 1 s of that, however many threads you add. The output is byte-identical whatever you choose.

Small jobs stay single-threaded on purpose — handing the work to other threads would cost more than the alignments it saves. Choosing an explicit thread count overrides that. A page opened straight from disk (file://) cannot start Workers at all and falls back to one thread automatically; the Log tab always reports how many were really used.

  • Scale. This is intended for tens to a few hundred fragments, not for whole-genome shotgun data. Large inputs are accepted but the pairwise stage grows with the number of candidate pairs.
  • No paired-end or mate-pair information. Repeats longer than your fragments cannot be resolved by any method without long-range links, and the tool stops rather than guessing.
  • Contig orientation is arbitrary. With no reference, a contig may come out reverse-complemented relative to what you expect; that is not an error. An R in one orientation is a Y in the other.
  • Nothing is uploaded. The whole assembly runs in your browser.
  • Everything is one contig and it should not be — raise Min. identity, raise Min. overlap, or lower Max. unmatched overhang.
  • Nothing joins — check the Rejected overlaps table for the reason. Usually the fragments overlap by less than Min. overlap, or are more divergent than Min. identity allows. Fragments shorter than the seed word length k can never be joined.
  • Expected polymorphism is missing — check the coverage at that column in the Alignment view. Two reads cannot establish a minority allele; the thresholds and the statistical test both need enough depth.

TotalRepeats

Applies to: TotalRepeats — repeat identification and masking

Rapid de novo identification, masking, visualization, and clustering of all repetitive sequences at genomic scale. Detects direct and inverted repeats, microsatellites (SSRs), telomeric sequences, and complex higher-order repeat structures.

  • Direct repeats: tandemly arranged identical or near-identical sequences on the same strand.
  • Inverted repeats: palindromic sequences that can form hairpin/stem-loop secondary structures.
  • Microsatellites (SSRs): short tandem repeat units (2–6 nt motifs) repeated in tandem arrays.
  • Telomeric sequences: species-specific hexanucleotide repeat units at chromosome ends.
  • Higher-order repeats: complex arrays of repeated blocks built from other repeat types.

  1. Paste or upload one or more sequences in FASTA format.
  2. Set the k-mer size and Min. repeat length, choose the Masking of repeats mode, and tick Sensitive detection or Family consensus & copy divergence if you need them. There is no repeat-type selector and no copy-number threshold.
  3. Click Analysis. Results are reported per sequence with coordinates and consensus repeat units.
  4. Choose the Masking of repeats mode — Soft (lowercase, keeps the bases; the default) or Hard (replace with N or X, which is what BLAST and primer design expect) — then take the repeat-masked FASTA from the Masked sequences tab with Save Mask and paste it into the PCR / LAMP / KASP inputs. The Report tab exports the same findings as TSV, GFF3, BED or a consensus FASTA library.

Repeat-rich regions are a major source of primer design failures. The recommended workflow is:

  1. Run TotalRepeats on your target sequence to identify repeat coordinates.
  2. Export the masked sequence with Save Mask — repeat regions in lower case under the default Soft masking, or replaced by N or X if you pick a Hard mode — or manually add /…/ exclusion markup around identified repeat coordinates.
  3. Paste the masked/marked sequence into the PCR, KASP, LAMP, or Assembly tool for repeat-aware primer design.

The Picture tab draws a map of the detected repeat blocks and clusters along the sequence on a zoomable canvas.

  • Zoom: scroll the mouse wheel over the map (zoom centres on the cursor).
  • Pan: click and drag left or right.
  • Double-click: recentres and resets the view to the default zoom, showing the whole sequence from the start — use it whenever the map has drifted off-centre.

Two buttons below the map save the picture:

  • Save Image — exports the current view as a PNG raster image (map.png).
  • Save SVG — exports the current view as a scalable SVG vector file (map.svg) that stays sharp at any zoom and can be edited in vector-graphics software.

Both buttons capture the current zoom and pan, so double-click first to reset the view if you want the entire map in the saved file.

The browser-based tool is suitable for sequences up to 100 Mb. For whole-genome analyses, use the command-line TotalRepeats Java application, which supports unlimited input size.


PCR / LAMP Reaction Mix Calculator

Applies to: PCR, LAMP or any reaction setup calculator

Calculates the exact volume of each reagent to add when preparing PCR, qPCR, or LAMP reaction mixtures. Choose a preset polymerase/kit to auto-fill typical stock and final concentrations, then scale to any number of reactions and reaction volume.

  1. Select your polymerase or kit from the drop-down menu. Default values for buffer, MgCl₂/MgSO₄, dNTP, and polymerase concentrations are loaded automatically.
  2. Enter the number of reactions and reaction volume (µl). The total master-mix volume is computed automatically.
  3. Adjust stock and final concentrations for any component (primers, probes, DMSO, ROX, SYBR Green II, etc.) to match your specific reagents.
  4. The output table updates instantly and lists the volume (µl) of each component to pipette.

PCR / High-fidelity
  • Standard Taq DNA Polymerase
  • Phusion DNA Polymerase
  • Phire Hot Start II DNA Polymerase
  • Deep VentR DNA Polymerase
  • KAPA HiFi DNA Polymerase
  • Herculase II Fusion DNA Polymerase
  • Pfu DNA Polymerase
Long / Isothermal
  • Long PCR (Taq + Pfu blend)
  • LongAmp Taq DNA Polymerase
  • LAMP (Bst 2.0 Polymerase — isothermal buffer with MgSO₄, Betaine support)

All reagent labels and concentrations are editable. Available slots include:

  • PCR buffer (label editable)
  • MgCl₂ or MgSO₄ (mM)
  • dNTP mix (mM)
  • Forward and reverse primers (µM) — or FIP/BIP for LAMP
  • Up to 4 additional probes or primers (e.g., F3/B3, loop primers for LAMP)
  • An extra component slot with a custom unit label (e.g., Betaine, enzyme enhancer)
  • DMSO (%), ROX reference dye (×), SYBR Green II (×)
  • Primary polymerase (U/µl) and optional secondary Pfu polymerase (U/µl)
  • DNA template volume (µl)

  • Set a final concentration of 0 for any component you do not need — it will be excluded from the output.
  • Include a 5–10% overage by entering a slightly higher number of reactions (e.g., 11 instead of 10) to account for pipetting losses.
  • For LAMP, the preset loads the full 6-primer set (FIP, BIP, F3, B3, FLOOP, BLOOP) with recommended concentrations and isothermal buffer.
  • Export the result table with Ctrl+A → Ctrl+C → paste into Excel or your lab notebook.

Universal Two-Solution Mixing and Dilution Calculator

Applies to: Universal Two-Solution Mixing and Dilution Calculator

A general-purpose laboratory calculator for mixing two solutions of known concentrations. Provide any 4 of the 6 variables and the remaining 2 are computed automatically. Works with any consistent unit system (M, mM, %, mass fraction; volumes in L, mL, or µl).

Solution 1 (Stock)
  • C₁ — concentration of the stock solution
  • V₁ — volume of stock to use
Solution 2 (Solvent/Diluent)
  • C₂ — concentration of the solvent (use 0 for pure water/buffer)
  • V₂ — volume of solvent to add
Final (mixed) solution
  • Cf — desired final concentration
  • Vf — desired final volume (= V₁ + V₂)

0 is a valid entry (e.g., C₂ = 0 for dilution with pure water). Leave a field empty only when it is truly unknown.

#TaskUnknown(s)
1Mix 400 mL of 25% with 100 mL of 15% — find final concentrationCf
2Mix pH 9.0 and pH 7.0 to reach pH 7.6 in 100 mL (linear approx.)V₁, V₂
3Dilute 100 mL of 0.5 M stock with 200 mL water — find concentrationCf
4Blend 10% and 5% to get 30 gal of 7% solutionV₁, V₂
5Add 10% to 40 gal of 35% to reach 20%V₂, Vf
6Mix 60% and 20% alcohol to make 20 gal of 30%V₁, V₂
7Blend 90% gold with unknown alloy for 30 oz of 80% goldC₂, V₂

  • Units must be consistent across all concentration fields (all in mM, all in %, etc.) and across all volume fields (all in mL, µl, etc.).
  • pH mixing uses a linear approximation (C₁·V₁ + C₂·V₂ = Cf·Vf), valid only for rough educational estimates; true pH requires accounting for buffer capacity and activity coefficients.
  • The tool assumes ideal mixing (volumes are additive). In practice, some solutions (e.g., concentrated sulfuric acid) show significant volume changes on mixing.
  • The formula applied is the standard mixing equation: C₁V₁ + C₂V₂ = CfVf.

Exporting results

  • Most design tools now carry Copy and Download TSV buttons above each result area — use those in preference to the clipboard. Sequence Utilities has Copy and a Download that names the file after the format it holds; TotalRepeats has Download TSV / GFF3 / BED / consensus (FASTA) on the Report tab, Save Mask on the Masked sequences tab and Save Image / Save SVG on the picture tab.
  • For tables in <textarea> outputs: click inside, then Ctrl+A (select all) → Ctrl+C (copy) → paste into Excel/Sheets.
  • If a tool produces multiple outputs (e.g., primer list and primer pairs), switch tabs before copying.
  • When pasting into spreadsheets, use "Split text to columns" if a fixed delimiter (space or tab) is used.
  • For PrimersList's Features tab, horizontal scrolling may be needed — paste with wrapping disabled (wrap="off" is set).

Troubleshooting

TSDR, the tree builder and Sequence Utilities check your input before they run. If something is wrong, a message appears next to the input, above the run button, rather than the tool failing silently:
  • Red — the run was stopped. Empty input, too little sequence, no nucleotide bases at all, a protein sequence where nucleotides are needed, unbalanced [ ] brackets, an empty [] bracket, a missing variant marker on TSDR, or fewer than three sequences on the tree builder. Fix what the message names and run again. The tree builder is the exception: it also accepts a Newick string, which the sequence check cannot judge, so it always computes and a red message there is a warning about the input rather than a refusal.
  • Amber — the run went ahead, but read this. Most often it lists characters that are not valid nucleotide codes and were therefore ignored. That matters because dropping them shifts every position after them, so check the result is what you expected.

  • Verify that you clicked the tool's run button (Generate, Design Primers, Design Probes, …) and that output is shown in the correct result tab.
  • Reduce constraints (widen Tm range by 1–2°C; allow slightly longer/shorter primers) and try again.
  • Confirm the sequence contains valid characters only (A/C/G/T and allowed degenerate codes); remove spaces or non-ASCII characters.
  • Open the browser console (F12 → Console) and check for JavaScript errors.

  • Reduce the probe distance parameter if the target region is short — a very large required gap can exhaust available candidates.
  • Confirm the input sequence is long enough: two non-overlapping 35–40 nt probes plus the required distance gap requires at least 75–80 nt of usable target sequence.
  • Remove or narrow any /.../ exclusion regions — they may be blocking too much of the target.
  • Verify you are not targeting the ITR region, which is excluded by design.
  • Probe length is fixed at 35–40 nt and one value is used for the whole run — anything outside that range is silently clamped. Try 35 nt rather than a longer setting, since shorter probes have more valid placements.

  • Confirm all final concentrations and the template volume are filled in correctly — any missing entry is treated as 0 and excluded.
  • The calculator works from the total premix volume — reaction volume × number of reactions, shown in the Total volume field. Water (or buffer) is added automatically to reach that total; if the components already exceed it, a negative water volume will be shown. With the default Number of reactions of 1 the premix and the single reaction are the same volume.
  • Check that stock concentrations are in the correct units — stock and final concentrations must use the same unit system (both in µM, or both in mM, etc.).

  • Confirm exactly 4 fields are filled and 2 are left empty. Providing 5 or 6 values over-constrains the system.
  • For pure dilution (adding water/buffer), set C₂ = 0, not blank — blank means "unknown".
  • If a calculated volume is negative, the combination of concentrations and volumes is physically impossible (e.g., the desired final concentration is higher than both stock concentrations).
  • Ensure all concentration values use the same units and all volume values use the same units.

If the page suddenly stops working after a server update (buttons do nothing, tabs don't switch, no results), the browser is likely using a cached JavaScript file.

Hard reload
  • Windows/Linux: Ctrl+F5 or Ctrl+Shift+R
  • macOS: Cmd+Shift+R
Chrome / Edge (strongest option)
  1. Press F12 to open DevTools
  2. Right-click the reload button → Empty cache and hard reload
Recommended server-side fix (prevents recurrence)
Add versioned URLs for scripts and CSS (cache-busting), e.g. ../js/panel.js?v=20260312. Update the version string after every deployment.

  • Confirm the ID is correct (e.g., NCBI nucleotide accession like A02710; rsIDs like rs4988235).
  • Network/firewall policies may block external API calls; use manual FASTA upload instead.
  • Try again later if the upstream service is rate-limiting or temporarily unavailable.

  • Enable Non-specific priming control and increase the minimum complexity threshold.
  • Run TotalRepeats on your target to identify repeat coordinates, then add /.../ exclusion markup around them.
  • For multiplex panels: tighten the primer Tm window.

  • Provide at least three fragments in assembly order — typically vector-left, insert(s), vector-right. With fewer, the tool prints “Gibson Assembly needs at least 3 fragments in assembly order” instead of primers.
  • Ensure sequences are in the correct assembly order in the FASTA file.
  • Verify that adjacent fragments share ≥20 bp overlapping sequence with a Tm ≥50°C for exonuclease processing.
  • If using vector sequences at the ends, confirm they contain the appropriate homology regions matching the adjacent insert ends.
  • Widen the Tm window (e.g., 58–64°C), or lower Min. length (nt) so shorter annealing regions qualify. There is no maximum-length setting — the primer is extended until it reaches the minimum Tm.

  • Try increasing the Max F2-B2 amplicon size (default 200 bp; try up to 350 bp).
  • Widen the Tm range slightly (e.g., 58–63°C) to allow more inner-primer candidates.
  • Disable Non-specific priming control temporarily to confirm a set can be generated on the current sequence, then re-enable it.
  • Use [ ... ] markup to direct the tool to the most unique subregion of a long target sequence.