Introduction & Overview
CHOZENLAB Sequence V2 is a professional-grade, privacy-first computational biology platform built for researchers, molecular biologists, students, and bioinformaticians. It enables comprehensive analysis of nucleotide (DNA/RNA) and protein sequences entirely within a web browser — no software installation required.
The platform is powered by Biopython — the gold standard open-source library for bioinformatics in Python — running server-side via a FastAPI REST backend. All analytical results are scientifically reproducible and mathematically grounded.
Quick Start Guide
Get your first sequence analyzed in under 60 seconds:
- 1Navigate to the Sequence AnalyzerGo to the homepage. You will see the main sequence input panel.
- 2Paste or Upload a SequencePaste a raw nucleotide or protein sequence, or upload a .fasta / .fa / .txt file. Multi-FASTA files are supported — a record selector will appear automatically.
- 3Select Sequence TypeLeave the type on "Auto" to let the platform detect whether the sequence is DNA, RNA, or Protein. You can override this manually.
- 4Click AnalyzeHit the "Analyze" button. Results appear across multiple tabs: Statistics, Sequence Viewer, ORF Explorer, Translation, and GC Profile.
- 5Export Your ResultsUse the Export menu to download JSON, TXT report, FASTA, or CSV formats depending on your downstream needs.
Sequence Input & FASTA Support
CHOZENLAB accepts sequences in two formats:
Raw Sequence (Plain Text)
Simply paste the nucleotide or amino acid sequence directly. Whitespace, newlines, and numbers (e.g., from GenBank formatting) are automatically stripped.
ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAGCGCATCGATCGATCGAT
FASTA Format (Single or Multi-Record)
FASTA headers are automatically parsed. For multi-FASTA files, a record selector appears at the top of the results so you can switch between sequences. Each record is analyzed independently.
>NC_001416.1 Enterobacteria phage lambda GGGCGGCGACCTCGCGGGTTTTCGCTATTTATGAAAATTTTCCGGTTTAAGGCGTTTCCGTTCTTCTTCG TATAATGTTTTAATCTTTTGTTTTGAACATTTAATCCTTTTTTTTTATTTCTCGTTTGAGGGTTGTGATG >NM_007298.3 BRCA1 mRNA ATGGATTTATCTGCTCTTCGCGTTGAAGAAGTACAAAATGTCATTAATGCTATGCAGAAAATCTTAG
Supported file extensions: .fasta, .fa, .fna, .faa, .txt
DNA Sequence Analysis
When a DNA sequence is detected, CHOZENLAB computes a full suite of structural and compositional metrics:
| Metric | Formula / Method | Significance |
|---|---|---|
| Sequence Length | Count of valid bases | Fundamental genomic metric |
| Base Counts (A, T, G, C) | Absolute frequency per nucleotide | Foundation for all downstream stats |
| GC Content (%) | (G + C) / Length × 100 | Genome stability, species identification, primer design |
| AT Content (%) | (A + T) / Length × 100 | Complement of GC content |
| GC Skew | (G − C) / (G + C) | Identifies leading/lagging replication strands; detects origin of replication |
| AT Skew | (A − T) / (A + T) | Strand asymmetry; useful for replication and mutational bias studies |
| Molecular Weight (Da) | Sum of average nucleotide monoisotopic masses | Physical mass estimate; relevant for synthesis and labeling |
| Reverse Complement | 5′→3′ complement of reversed sequence | Represents the antisense / template strand |
| Complement | A↔T, G↔C substitution (no reversal) | Coding complement without strand orientation change |
| Transcription | Replace T with U (DNA → mRNA) | Predicts the messenger RNA transcript sequence |
RNA Sequence Analysis
RNA sequences (containing Uracil instead of Thymine) are automatically detected and analyzed with RNA-specific metrics. CHOZENLAB uses Biopython's Bio.Seq to handle RNA alphabet natively.
| Metric | Description |
|---|---|
| Base Counts (A, U, G, C) | Frequency of all four RNA nucleotides |
| GC Content | (G + C) / Total × 100 — identical formula to DNA |
| GC Skew | (G − C) / (G + C) — identifies structural features in mRNA |
| AU Skew | (A − U) / (A + U) — RNA-specific asymmetry metric, analogous to AT skew |
| Molecular Weight | Calculated using Biopython's molecular_weight() with seq_type="RNA" |
| Reverse Complement | RNA complement using A↔U, G↔C rules |
| Back-Transcription | Converts RNA back to the encoding DNA sequence (U → T) |
| Translation | Translates RNA codons to protein sequence using the Standard Genetic Code |
Protein Sequence Analysis
Protein sequences are analyzed using Biopython's ProtParam module, which implements well-established physicochemical prediction algorithms.
Molecular Weight (Da)+
Isoelectric Point (pI)+
Instability Index+
Aromaticity (Lobry & Gautier)+
Aliphatic Index (Ikai)+
Amino Acid Composition+
Residue Group Counts+
Six-Frame Translation
Any double-stranded DNA sequence can be read in six possible reading frames: three on the forward (5′→3′) strand and three on the reverse complement (3′→5′ strand, read 5′→3′). CHOZENLAB automatically computes all six.
For each frame, the amino acid sequence is computed using Biopython's Seq.translate() with the Standard Genetic Code (NCBI Table 1). Start codons (AUG) are highlighted in green; stop codons (*) are highlighted in red.
The Translation tab shows all six frames simultaneously with codon-level color coding, nucleotide position coordinates, and per-frame copy buttons.
ORF Explorer
An Open Reading Frame (ORF) is a sequence of DNA that begins with a start codon and ends with a stop codon, uninterrupted in the same reading frame. CHOZENLAB's ORF Explorer automatically identifies all candidate ORFs from all six reading frames.
Detection Criteria
- Start Codon: ATG (methionine) — canonical initiation signal in all standard genetic codes.
- Stop Codon: TAA, TAG, or TGA — canonical termination signals.
- Minimum Length: ≥30 amino acids (90 nucleotides). Shorter ORFs are excluded to minimize false positives.
- All six frames are searched simultaneously on the original sequence and its reverse complement.
ORF Table Fields
| Field | Description |
|---|---|
| ORF ID | Sequential identifier (ORF-001, ORF-002...) ordered by position |
| Frame | Reading frame: +1, +2, +3 (forward) or −1, −2, −3 (reverse) |
| Strand | + (forward/sense) or − (reverse/antisense) |
| Start | 1-based nucleotide position of the ATG start codon |
| End | 1-based nucleotide position of the last nucleotide of the stop codon |
| Length (bp) | Total nucleotide length of the ORF including start and stop codons |
| Length (aa) | Number of amino acids encoded (excluding stop codon) |
| Start Codon | Always ATG for standard genetic code |
| Stop Codon | TAA, TAG, or TGA — the specific stop codon used |
| Protein Sequence | The translated amino acid sequence (visible in the detail panel) |
Scientific Disclaimer: These are candidate ORFs identified by computational structure only. Biological validation via wet lab experiments is required before any functional claims can be made.
GC Content Profile
For nucleotide sequences, CHOZENLAB generates a GC Content Profile — a sliding window plot showing how GC percentage varies along the length of the sequence. This is a standard tool in genomic analysis used to:
- Identify isochores (regions of uniform GC composition)
- Locate CpG islands (high GC regions often upstream of promoters)
- Detect genomic islands (horizontally transferred regions with atypical composition)
- Distinguish coding vs. non-coding regions (coding regions tend to have higher GC content)
- Identify origins of replication through GC skew sign changes
Pairwise Sequence Alignment
Available at /tools/alignment, the pairwise alignment tool compares two nucleotide or protein sequences using dynamic programming algorithms implemented by Biopython's PairwiseAligner.
Global Alignment (Needleman-Wunsch)
Forces an end-to-end alignment spanning the entire length of both sequences. Gaps are introduced as needed anywhere in either sequence to produce the optimal global score.
Local Alignment (Smith-Waterman)
Identifies the highest-scoring local region of similarity between two sequences, without requiring full-length coverage. Produces sub-sequence alignments.
Alignment Output Metrics
| Metric | Formula |
|---|---|
| Alignment Score | Raw score from the dynamic programming matrix |
| Identity (%) | Identical positions / alignment length × 100 |
| Mismatches | Positions where bases differ (non-gap) |
| Gaps | Total gap characters in the aligned output |
| Aligned Length | Total columns in the alignment including gaps |
Performance note: Alignment is limited to 10,000 bp per sequence for browser responsiveness. BLAST is recommended for larger sequences.
PCR Primer Analysis
Available at /tools/primer, the Primer Analysis tool validates forward and reverse PCR primers against a DNA template. It uses Biopython's MeltingTemp module for thermodynamic calculations.
Melting Temperature (Tm) Methods
A quick approximation: Tm = 2(A+T) + 4(G+C). Best for primers of 14–20 bp length. Fast but less accurate for GC-rich or long primers.
Uses experimentally derived ΔH and ΔS values for each dinucleotide pair: Tm = ΔH / (ΔS + R × ln[primer]) − 273.15. Far more accurate for any primer length.
Primer Design Rules Checked
- Length: Optimal range 18–25 bp. Outside this range triggers a warning.
- GC Content: Ideal 40–60%. Too high or low GC can cause non-specific binding or low binding affinity.
- 3′ GC Clamp: The last 1–3 bases at the 3′ end should be G or C to ensure stable polymerase priming.
- Template Binding: Verifies the primer sequence is present in the provided template. Checks both forward strand (forward primer) and reverse complement (reverse primer).
- Amplicon Length: Calculates the expected PCR product size from the binding positions of both primers.
Restriction Enzyme Mapping
Available at /tools/restriction, this tool uses Biopython's Bio.Restriction module backed by the REBASE (Restriction Enzyme Database) to find all cut sites of selected restriction endonucleases.
Common enzymes available include:
For each selected enzyme, the tool reports: recognition sequence, cut position notation, number of cut sites, and all positions (1-based) along the query sequence.
Exporting Results
All analysis results can be exported in multiple formats. Export buttons are available in the top-right of the results dashboard.
Complete machine-readable export of all computed fields. Ideal for downstream scripting, data pipelines, or custom visualizations. Contains every metric, ORF, codon table, reading frames, and alignment data.
A human-readable scientific report in plain text format. Includes sequence type, all statistics, ORF summary, primer results, and a full scientific disclaimer section. Suitable for sharing with collaborators or attaching to lab notebooks.
Exports the sequence along with enriched FASTA headers containing GC content, MW, ORF count, and analysis timestamp. Useful for submitting to external databases or archiving.
Amino acid composition table exported as comma-separated values. Directly importable into Excel, R, or Python/pandas for statistical analysis.
Biopython Code Reference
CHOZENLAB champions scientific reproducibility. Every calculation in this platform can be independently replicated using Biopython. Below is a complete reference of our core implementation.
pip install biopython
DNA: Base Analysis & Transcription
Bio.Seq docs ↗from Bio.Seq import Seq
dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")
# Base counts
a, t, g, c = dna.count('A'), dna.count('T'), dna.count('G'), dna.count('C')
gc = round((g + c) / len(dna) * 100, 2)
gc_skew = round((g - c) / (g + c), 4)
at_skew = round((a - t) / (a + t), 4)
print(f"GC%: {gc}")
print(f"GC Skew: {gc_skew} AT Skew: {at_skew}")
print("Complement:", dna.complement())
print("Reverse Complement:", dna.reverse_complement())
print("Transcription (→ mRNA):", dna.transcribe())Six-Frame Translation & ORF Detection
Bio.Seq docs ↗from Bio.Seq import Seq
dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")
revcomp = dna.reverse_complement()
frames = {}
for i in range(3):
frames[f"+{i+1}"] = dna[i:].translate(to_stop=False)
frames[f"-{i+1}"] = revcomp[i:].translate(to_stop=False)
# Find ORFs (start to stop, min 30 aa)
def find_orfs(seq_str, min_aa=30):
seq = Seq(seq_str)
orfs = []
for i in range(len(seq) - 2):
if seq[i:i+3] == "ATG":
remainder = seq[i:]
protein = remainder[:len(remainder) - len(remainder) % 3].translate()
stop = str(protein).find("*")
if stop >= min_aa:
orfs.append({"start": i+1, "end": i+(stop+1)*3, "protein": str(protein[:stop])})
return orfsPairwise Sequence Alignment
Bio.Align docs ↗from Bio import Align
aligner = Align.PairwiseAligner()
aligner.mode = 'global' # or 'local' for Smith-Waterman
aligner.match_score = 2
aligner.mismatch_score = -1
aligner.open_gap_score = -2
aligner.extend_gap_score = -0.5
seq1 = "GCTAGCTACGATCGAT"
seq2 = "GCAGCTACGTTCGAT"
alignments = aligner.align(seq1, seq2)
best = alignments[0]
print(best)
print(f"Score: {best.score}")
print(f"Identity: {best.counts().identities}/{best.length} = {best.counts().identities/best.length*100:.1f}%")Restriction Enzyme Mapping
Bio.Restriction docs ↗from Bio.Seq import Seq
from Bio.Restriction import RestrictionBatch, Analysis
dna = Seq("GAATTCGGATCCAAGCTTCATATGGTACCC")
enzymes = RestrictionBatch(["EcoRI", "BamHI", "HindIII", "NdeI", "KpnI"])
analysis = Analysis(enzymes, dna, linear=True)
results = analysis.full()
for enzyme, sites in results.items():
if sites:
print(f"{enzyme}: cuts at positions {sites}")PCR Primer Melting Temperature (Tm)
MeltingTemp docs ↗from Bio.SeqUtils import MeltingTemp as mt
primer = "GCTAGCTACGATCGAT"
# Wallace Rule (simple approximation)
tm_wallace = mt.Tm_Wallace(primer)
# Nearest-Neighbor thermodynamics (more accurate)
tm_nn = mt.Tm_NN(primer)
# Primer stats
gc = (primer.count('G') + primer.count('C')) / len(primer) * 100
has_clamp = primer[-1] in 'GC' or primer[-2] in 'GC' or primer[-3] in 'GC'
print(f"Length: {len(primer)} bp")
print(f"GC Content: {gc:.1f}%")
print(f"Tm (Wallace): {tm_wallace:.2f} °C")
print(f"Tm (NN): {tm_nn:.2f} °C")
print(f"3' GC Clamp: {'YES' if has_clamp else 'NO'}")Protein Physicochemical Properties
ProtParam docs ↗from Bio.SeqUtils.ProtParam import ProteinAnalysis
protein = "MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRILPDGTVDGTP"
analysis = ProteinAnalysis(protein)
print(f"Length: {len(protein)} aa")
print(f"MW: {analysis.molecular_weight():.2f} Da")
print(f"pI: {analysis.isoelectric_point():.2f}")
print(f"Instability: {analysis.instability_index():.2f}")
print(f"Aromaticity: {analysis.aromaticity():.4f}")
print(f"Aliphatic Index: {analysis.aliphatic_index():.2f}")
# Amino acid composition
composition = analysis.get_amino_acids_percent()
print("\nTop 3 residues:")
for aa, pct in sorted(composition.items(), key=lambda x: -x[1])[:3]:
print(f" {aa}: {pct*100:.1f}%")GC Content Sliding Window Profile
Bio.SeqUtils docs ↗from Bio.SeqUtils import gc_fraction
def gc_profile(sequence, window=100, step=10):
"""Compute GC% in a sliding window across the sequence."""
results = []
for i in range(0, len(sequence) - window + 1, step):
window_seq = sequence[i:i+window]
gc = gc_fraction(window_seq) * 100
results.append({"position": i + window // 2, "gc_percent": round(gc, 2)})
return results
dna = "ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG" * 10
profile = gc_profile(dna)
for point in profile[:5]:
print(f"Position {point['position']}: GC = {point['gc_percent']:.2f}%")RNA Analysis
Bio.Seq docs ↗from Bio.Seq import Seq
from Bio.SeqUtils import molecular_weight
rna = Seq("AUGGCCAUUGUAAUGGGCCGCUGAAAGGGUGCCCGAUAG")
# Base counts
counts = {b: rna.count(b) for b in 'AUGC'}
gc = round((counts['G'] + counts['C']) / len(rna) * 100, 2)
gc_skew = round((counts['G'] - counts['C']) / (counts['G'] + counts['C']), 4)
au_skew = round((counts['A'] - counts['U']) / (counts['A'] + counts['U']), 4)
# MW of RNA
mw = molecular_weight(str(rna), seq_type='RNA')
# Back-transcribe to DNA
dna = rna.back_transcribe()
print(f"GC%: {gc} GC Skew: {gc_skew} AU Skew: {au_skew}")
print(f"MW: {mw:.2f} Da")
print(f"Back-transcribed DNA: {dna}")
print(f"Translation: {rna.translate()}")Data Privacy & Security
All sequences processed by CHOZENLAB Sequence V2 are handled ephemerally. When you submit a sequence for analysis, it is transmitted securely via HTTPS to our FastAPI backend, processed entirely in RAM, and the results are immediately returned to your browser.
Your raw sequences are never written to any database or persistent storage. There are no user accounts, no session tracking, and no long-term retention of biological data.
Frequently Asked Questions
What file formats does CHOZENLAB accept?+
Is there a sequence length limit?+
Are ORF predictions biologically validated?+
Which genetic code does the translation use?+
Can I analyze virus or bacterial genomes?+
Is CHOZENLAB Sequence free to use?+
How do I cite CHOZENLAB Sequence in a publication?+
Can I reproduce the results in my own Python environment?+
Ready to analyze your sequence?
Paste a DNA, RNA, or protein sequence and get a full report in seconds — no login required.
Open Sequence Analyzer →