CHOZENLAB — Biomedical Technologies & Life Sciences

CHOZENLAB Sequence V2
Documentation

A comprehensive reference for every feature in CHOZENLAB Sequence V2 — the browser-based bioinformatics suite for DNA, RNA, and protein sequence analysis. All computations are performed server-side using Biopython and FastAPI, with a Next.js frontend for instant, interactive results.

DNA AnalysisRNA AnalysisProtein AnalysisORF DetectionSequence AlignmentPrimer DesignRestriction MappingFASTA Support

Introduction & Overview

CHOZENLAB Sequence V2 is a professional-grade, privacy-first computational biology platform built for researchers, molecular biologists, students, and bioinformaticians. It enables comprehensive analysis of nucleotide (DNA/RNA) and protein sequences entirely within a web browser — no software installation required.

The platform is powered by Biopython — the gold standard open-source library for bioinformatics in Python — running server-side via a FastAPI REST backend. All analytical results are scientifically reproducible and mathematically grounded.

Max Sequence Length
500,000 bp
General structural analysis
Alignment Limit
10,000 bp
Pairwise alignment performance
FASTA Multi-Record
Unlimited
Switch between records freely

Quick Start Guide

Get your first sequence analyzed in under 60 seconds:

  1. 1
    Navigate to the Sequence Analyzer
    Go to the homepage. You will see the main sequence input panel.
  2. 2
    Paste or Upload a Sequence
    Paste a raw nucleotide or protein sequence, or upload a .fasta / .fa / .txt file. Multi-FASTA files are supported — a record selector will appear automatically.
  3. 3
    Select Sequence Type
    Leave the type on "Auto" to let the platform detect whether the sequence is DNA, RNA, or Protein. You can override this manually.
  4. 4
    Click Analyze
    Hit the "Analyze" button. Results appear across multiple tabs: Statistics, Sequence Viewer, ORF Explorer, Translation, and GC Profile.
  5. 5
    Export Your Results
    Use the Export menu to download JSON, TXT report, FASTA, or CSV formats depending on your downstream needs.

Sequence Input & FASTA Support

CHOZENLAB accepts sequences in two formats:

Raw Sequence (Plain Text)

Simply paste the nucleotide or amino acid sequence directly. Whitespace, newlines, and numbers (e.g., from GenBank formatting) are automatically stripped.

ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAGCGCATCGATCGATCGAT

FASTA Format (Single or Multi-Record)

FASTA headers are automatically parsed. For multi-FASTA files, a record selector appears at the top of the results so you can switch between sequences. Each record is analyzed independently.

>NC_001416.1 Enterobacteria phage lambda
GGGCGGCGACCTCGCGGGTTTTCGCTATTTATGAAAATTTTCCGGTTTAAGGCGTTTCCGTTCTTCTTCG
TATAATGTTTTAATCTTTTGTTTTGAACATTTAATCCTTTTTTTTTATTTCTCGTTTGAGGGTTGTGATG

>NM_007298.3 BRCA1 mRNA
ATGGATTTATCTGCTCTTCGCGTTGAAGAAGTACAAAATGTCATTAATGCTATGCAGAAAATCTTAG

Supported file extensions: .fasta, .fa, .fna, .faa, .txt

DNA Sequence Analysis

When a DNA sequence is detected, CHOZENLAB computes a full suite of structural and compositional metrics:

MetricFormula / MethodSignificance
Sequence LengthCount of valid basesFundamental genomic metric
Base Counts (A, T, G, C)Absolute frequency per nucleotideFoundation for all downstream stats
GC Content (%)(G + C) / Length × 100Genome stability, species identification, primer design
AT Content (%)(A + T) / Length × 100Complement of GC content
GC Skew(G − C) / (G + C)Identifies leading/lagging replication strands; detects origin of replication
AT Skew(A − T) / (A + T)Strand asymmetry; useful for replication and mutational bias studies
Molecular Weight (Da)Sum of average nucleotide monoisotopic massesPhysical mass estimate; relevant for synthesis and labeling
Reverse Complement5′→3′ complement of reversed sequenceRepresents the antisense / template strand
ComplementA↔T, G↔C substitution (no reversal)Coding complement without strand orientation change
TranscriptionReplace T with U (DNA → mRNA)Predicts the messenger RNA transcript sequence
GC Skew Interpretation: A positive skew (G > C) is typically found on the leading strand. Switching between positive and negative GC skew along a genome often pinpoints the origin of replication (oriC).

RNA Sequence Analysis

RNA sequences (containing Uracil instead of Thymine) are automatically detected and analyzed with RNA-specific metrics. CHOZENLAB uses Biopython's Bio.Seq to handle RNA alphabet natively.

MetricDescription
Base Counts (A, U, G, C)Frequency of all four RNA nucleotides
GC Content(G + C) / Total × 100 — identical formula to DNA
GC Skew(G − C) / (G + C) — identifies structural features in mRNA
AU Skew(A − U) / (A + U) — RNA-specific asymmetry metric, analogous to AT skew
Molecular WeightCalculated using Biopython's molecular_weight() with seq_type="RNA"
Reverse ComplementRNA complement using A↔U, G↔C rules
Back-TranscriptionConverts RNA back to the encoding DNA sequence (U → T)
TranslationTranslates RNA codons to protein sequence using the Standard Genetic Code

Protein Sequence Analysis

Protein sequences are analyzed using Biopython's ProtParam module, which implements well-established physicochemical prediction algorithms.

Molecular Weight (Da)+
Calculates the average molecular weight by summing the monoisotopic masses of all amino acid residues and subtracting water (H₂O) for each peptide bond. Expressed in Daltons (Da).
Isoelectric Point (pI)+
The pH at which the total net charge of the protein is zero. Calculated by iteratively finding the pH where the sum of charge contributions from all basic (R, K, H) and acidic (D, E, C, Y, terminal residues) groups equals zero. Critical for gel electrophoresis and protein purification strategies.
Instability Index+
Predicts whether the protein is stable in a test tube. Based on the statistical analysis of 12 amino acid dipeptides (DIWV method). A value below 40 predicts a stable protein; above 40 predicts instability.
Aromaticity (Lobry & Gautier)+
The fraction of Phenylalanine (F), Tyrosine (Y), and Tryptophan (W) residues. Related to UV-absorption at 280 nm. Formula: (F + Y + W) / Length.
Aliphatic Index (Ikai)+
A measure of the relative volume occupied by aliphatic side chains — Alanine (A), Valine (V), Isoleucine (I), and Leucine (L). A higher aliphatic index indicates greater thermostability. Formula: AI = (A%) + 2.9(V%) + 3.9(I% + L%).
Amino Acid Composition+
A complete breakdown of all 20 standard amino acids with absolute counts and percentage frequencies. Displayed as a visual bar chart sorted by abundance. Full names (e.g., Alanine, Arginine) are shown alongside single-letter codes.
Residue Group Counts+
Counts of Basic residues (K, R, H), Acidic residues (D, E), total Charged residues, and Hydrophobic residues (A, V, I, L, M, F, W, P). Useful for predicting membrane interactions, solubility, and charge distribution.

Six-Frame Translation

Any double-stranded DNA sequence can be read in six possible reading frames: three on the forward (5′→3′) strand and three on the reverse complement (3′→5′ strand, read 5′→3′). CHOZENLAB automatically computes all six.

+1 (Forward)
+2 (Forward)
+3 (Forward)
−1 (Reverse)
−2 (Reverse)
−3 (Reverse)

For each frame, the amino acid sequence is computed using Biopython's Seq.translate() with the Standard Genetic Code (NCBI Table 1). Start codons (AUG) are highlighted in green; stop codons (*) are highlighted in red.

The Translation tab shows all six frames simultaneously with codon-level color coding, nucleotide position coordinates, and per-frame copy buttons.

Scientific Note: Translating all frames is the first step in ab initio gene prediction. However, not all translated frames are biologically meaningful — use the ORF Explorer to filter structurally valid candidates.

ORF Explorer

An Open Reading Frame (ORF) is a sequence of DNA that begins with a start codon and ends with a stop codon, uninterrupted in the same reading frame. CHOZENLAB's ORF Explorer automatically identifies all candidate ORFs from all six reading frames.

Detection Criteria

  • Start Codon: ATG (methionine) — canonical initiation signal in all standard genetic codes.
  • Stop Codon: TAA, TAG, or TGA — canonical termination signals.
  • Minimum Length: ≥30 amino acids (90 nucleotides). Shorter ORFs are excluded to minimize false positives.
  • All six frames are searched simultaneously on the original sequence and its reverse complement.

ORF Table Fields

FieldDescription
ORF IDSequential identifier (ORF-001, ORF-002...) ordered by position
FrameReading frame: +1, +2, +3 (forward) or −1, −2, −3 (reverse)
Strand+ (forward/sense) or − (reverse/antisense)
Start1-based nucleotide position of the ATG start codon
End1-based nucleotide position of the last nucleotide of the stop codon
Length (bp)Total nucleotide length of the ORF including start and stop codons
Length (aa)Number of amino acids encoded (excluding stop codon)
Start CodonAlways ATG for standard genetic code
Stop CodonTAA, TAG, or TGA — the specific stop codon used
Protein SequenceThe translated amino acid sequence (visible in the detail panel)
ORF Detail Panel: Click any ORF in the table to open a detailed panel showing the full nucleotide sequence (with reverse complement extraction for − strand ORFs), codon triplet visualization with start/stop highlighting, one-click copy buttons, and individual FASTA export.

Scientific Disclaimer: These are candidate ORFs identified by computational structure only. Biological validation via wet lab experiments is required before any functional claims can be made.

GC Content Profile

For nucleotide sequences, CHOZENLAB generates a GC Content Profile — a sliding window plot showing how GC percentage varies along the length of the sequence. This is a standard tool in genomic analysis used to:

  • Identify isochores (regions of uniform GC composition)
  • Locate CpG islands (high GC regions often upstream of promoters)
  • Detect genomic islands (horizontally transferred regions with atypical composition)
  • Distinguish coding vs. non-coding regions (coding regions tend to have higher GC content)
  • Identify origins of replication through GC skew sign changes
Technical Parameters
Window Size: 100 bp (default)
Step Size: 10 bp
Algorithm: Sliding window average
Backend Module: /api/gc-profile

Pairwise Sequence Alignment

Available at /tools/alignment, the pairwise alignment tool compares two nucleotide or protein sequences using dynamic programming algorithms implemented by Biopython's PairwiseAligner.

Global Alignment (Needleman-Wunsch)

Forces an end-to-end alignment spanning the entire length of both sequences. Gaps are introduced as needed anywhere in either sequence to produce the optimal global score.

Best for: Comparing full-length homologous sequences (e.g., orthologous genes from different species)
Scoring: Match +2, Mismatch −1, Gap open −2, Gap extend −0.5

Local Alignment (Smith-Waterman)

Identifies the highest-scoring local region of similarity between two sequences, without requiring full-length coverage. Produces sub-sequence alignments.

Best for: Detecting conserved domains or motifs within divergent sequences
Scoring: Match +2, Mismatch −1, Gap open −2, Gap extend −0.5

Alignment Output Metrics

MetricFormula
Alignment ScoreRaw score from the dynamic programming matrix
Identity (%)Identical positions / alignment length × 100
MismatchesPositions where bases differ (non-gap)
GapsTotal gap characters in the aligned output
Aligned LengthTotal columns in the alignment including gaps

Performance note: Alignment is limited to 10,000 bp per sequence for browser responsiveness. BLAST is recommended for larger sequences.

PCR Primer Analysis

Available at /tools/primer, the Primer Analysis tool validates forward and reverse PCR primers against a DNA template. It uses Biopython's MeltingTemp module for thermodynamic calculations.

Melting Temperature (Tm) Methods

Wallace Rule (Basic)

A quick approximation: Tm = 2(A+T) + 4(G+C). Best for primers of 14–20 bp length. Fast but less accurate for GC-rich or long primers.

Nearest-Neighbor (Thermodynamic)

Uses experimentally derived ΔH and ΔS values for each dinucleotide pair: Tm = ΔH / (ΔS + R × ln[primer]) − 273.15. Far more accurate for any primer length.

Primer Design Rules Checked

  • Length: Optimal range 18–25 bp. Outside this range triggers a warning.
  • GC Content: Ideal 40–60%. Too high or low GC can cause non-specific binding or low binding affinity.
  • 3′ GC Clamp: The last 1–3 bases at the 3′ end should be G or C to ensure stable polymerase priming.
  • Template Binding: Verifies the primer sequence is present in the provided template. Checks both forward strand (forward primer) and reverse complement (reverse primer).
  • Amplicon Length: Calculates the expected PCR product size from the binding positions of both primers.

Restriction Enzyme Mapping

Available at /tools/restriction, this tool uses Biopython's Bio.Restriction module backed by the REBASE (Restriction Enzyme Database) to find all cut sites of selected restriction endonucleases.

Common enzymes available include:

EcoRI (G↓AATTC)BamHI (G↓GATCC)HindIII (A↓AGCTT)NotI (GC↓GGCCGC)XhoI (C↓TCGAG)SalI (G↓TCGAC)PstI (CTGCA↓G)SmaI (CCC↓GGG)KpnI (GGTAC↓C)

For each selected enzyme, the tool reports: recognition sequence, cut position notation, number of cut sites, and all positions (1-based) along the query sequence.

Exporting Results

All analysis results can be exported in multiple formats. Export buttons are available in the top-right of the results dashboard.

{}JSON

Complete machine-readable export of all computed fields. Ideal for downstream scripting, data pipelines, or custom visualizations. Contains every metric, ORF, codon table, reading frames, and alignment data.

TXTTXT Report

A human-readable scientific report in plain text format. Includes sequence type, all statistics, ORF summary, primer results, and a full scientific disclaimer section. Suitable for sharing with collaborators or attaching to lab notebooks.

>FASTA

Exports the sequence along with enriched FASTA headers containing GC content, MW, ORF count, and analysis timestamp. Useful for submitting to external databases or archiving.

,CSV

Amino acid composition table exported as comma-separated values. Directly importable into Excel, R, or Python/pandas for statistical analysis.

Biopython Code Reference

CHOZENLAB champions scientific reproducibility. Every calculation in this platform can be independently replicated using Biopython. Below is a complete reference of our core implementation.

Installation
pip install biopython

DNA: Base Analysis & Transcription

Bio.Seq docs ↗
from Bio.Seq import Seq

dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")

# Base counts
a, t, g, c = dna.count('A'), dna.count('T'), dna.count('G'), dna.count('C')
gc = round((g + c) / len(dna) * 100, 2)
gc_skew = round((g - c) / (g + c), 4)
at_skew = round((a - t) / (a + t), 4)

print(f"GC%: {gc}")
print(f"GC Skew: {gc_skew}  AT Skew: {at_skew}")
print("Complement:", dna.complement())
print("Reverse Complement:", dna.reverse_complement())
print("Transcription (→ mRNA):", dna.transcribe())

Six-Frame Translation & ORF Detection

Bio.Seq docs ↗
from Bio.Seq import Seq

dna = Seq("ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG")
revcomp = dna.reverse_complement()

frames = {}
for i in range(3):
    frames[f"+{i+1}"] = dna[i:].translate(to_stop=False)
    frames[f"-{i+1}"] = revcomp[i:].translate(to_stop=False)

# Find ORFs (start to stop, min 30 aa)
def find_orfs(seq_str, min_aa=30):
    seq = Seq(seq_str)
    orfs = []
    for i in range(len(seq) - 2):
        if seq[i:i+3] == "ATG":
            remainder = seq[i:]
            protein = remainder[:len(remainder) - len(remainder) % 3].translate()
            stop = str(protein).find("*")
            if stop >= min_aa:
                orfs.append({"start": i+1, "end": i+(stop+1)*3, "protein": str(protein[:stop])})
    return orfs

Pairwise Sequence Alignment

Bio.Align docs ↗
from Bio import Align

aligner = Align.PairwiseAligner()
aligner.mode = 'global'  # or 'local' for Smith-Waterman
aligner.match_score = 2
aligner.mismatch_score = -1
aligner.open_gap_score = -2
aligner.extend_gap_score = -0.5

seq1 = "GCTAGCTACGATCGAT"
seq2 = "GCAGCTACGTTCGAT"

alignments = aligner.align(seq1, seq2)
best = alignments[0]

print(best)
print(f"Score: {best.score}")
print(f"Identity: {best.counts().identities}/{best.length} = {best.counts().identities/best.length*100:.1f}%")

Restriction Enzyme Mapping

Bio.Restriction docs ↗
from Bio.Seq import Seq
from Bio.Restriction import RestrictionBatch, Analysis

dna = Seq("GAATTCGGATCCAAGCTTCATATGGTACCC")
enzymes = RestrictionBatch(["EcoRI", "BamHI", "HindIII", "NdeI", "KpnI"])

analysis = Analysis(enzymes, dna, linear=True)
results = analysis.full()

for enzyme, sites in results.items():
    if sites:
        print(f"{enzyme}: cuts at positions {sites}")

PCR Primer Melting Temperature (Tm)

MeltingTemp docs ↗
from Bio.SeqUtils import MeltingTemp as mt

primer = "GCTAGCTACGATCGAT"

# Wallace Rule (simple approximation)
tm_wallace = mt.Tm_Wallace(primer)

# Nearest-Neighbor thermodynamics (more accurate)
tm_nn = mt.Tm_NN(primer)

# Primer stats
gc = (primer.count('G') + primer.count('C')) / len(primer) * 100
has_clamp = primer[-1] in 'GC' or primer[-2] in 'GC' or primer[-3] in 'GC'

print(f"Length: {len(primer)} bp")
print(f"GC Content: {gc:.1f}%")
print(f"Tm (Wallace): {tm_wallace:.2f} °C")
print(f"Tm (NN):      {tm_nn:.2f} °C")
print(f"3' GC Clamp: {'YES' if has_clamp else 'NO'}")

Protein Physicochemical Properties

ProtParam docs ↗
from Bio.SeqUtils.ProtParam import ProteinAnalysis

protein = "MAEGEITTFTALTEKFNLPPGNYKKPKLLYCSNGGHFLRILPDGTVDGTP"
analysis = ProteinAnalysis(protein)

print(f"Length:           {len(protein)} aa")
print(f"MW:               {analysis.molecular_weight():.2f} Da")
print(f"pI:               {analysis.isoelectric_point():.2f}")
print(f"Instability:      {analysis.instability_index():.2f}")
print(f"Aromaticity:      {analysis.aromaticity():.4f}")
print(f"Aliphatic Index:  {analysis.aliphatic_index():.2f}")

# Amino acid composition
composition = analysis.get_amino_acids_percent()
print("\nTop 3 residues:")
for aa, pct in sorted(composition.items(), key=lambda x: -x[1])[:3]:
    print(f"  {aa}: {pct*100:.1f}%")

GC Content Sliding Window Profile

Bio.SeqUtils docs ↗
from Bio.SeqUtils import gc_fraction

def gc_profile(sequence, window=100, step=10):
    """Compute GC% in a sliding window across the sequence."""
    results = []
    for i in range(0, len(sequence) - window + 1, step):
        window_seq = sequence[i:i+window]
        gc = gc_fraction(window_seq) * 100
        results.append({"position": i + window // 2, "gc_percent": round(gc, 2)})
    return results

dna = "ATGGCCATTGTAATGGGCCGCTGAAAGGGTGCCCGATAG" * 10
profile = gc_profile(dna)

for point in profile[:5]:
    print(f"Position {point['position']}: GC = {point['gc_percent']:.2f}%")

RNA Analysis

Bio.Seq docs ↗
from Bio.Seq import Seq
from Bio.SeqUtils import molecular_weight

rna = Seq("AUGGCCAUUGUAAUGGGCCGCUGAAAGGGUGCCCGAUAG")

# Base counts
counts = {b: rna.count(b) for b in 'AUGC'}
gc = round((counts['G'] + counts['C']) / len(rna) * 100, 2)
gc_skew = round((counts['G'] - counts['C']) / (counts['G'] + counts['C']), 4)
au_skew = round((counts['A'] - counts['U']) / (counts['A'] + counts['U']), 4)

# MW of RNA
mw = molecular_weight(str(rna), seq_type='RNA')

# Back-transcribe to DNA
dna = rna.back_transcribe()

print(f"GC%: {gc}  GC Skew: {gc_skew}  AU Skew: {au_skew}")
print(f"MW: {mw:.2f} Da")
print(f"Back-transcribed DNA: {dna}")
print(f"Translation: {rna.translate()}")

Data Privacy & Security

All sequences processed by CHOZENLAB Sequence V2 are handled ephemerally. When you submit a sequence for analysis, it is transmitted securely via HTTPS to our FastAPI backend, processed entirely in RAM, and the results are immediately returned to your browser.

Your raw sequences are never written to any database or persistent storage. There are no user accounts, no session tracking, and no long-term retention of biological data.

No Sequence Persistence
Sequences are processed in memory only and immediately discarded after the response.
No User Tracking
No personal data, cookies, or identifiers are stored. No login required.
Local Export Only
Use JSON, TXT, FASTA, or CSV exports to save results locally on your own machine.
Open Methodology
All algorithms are documented here and reproducible with Biopython — no black boxes.

Frequently Asked Questions

What file formats does CHOZENLAB accept?+
Raw sequences (plain text) and FASTA files (.fasta, .fa, .fna, .faa, .txt). FASTQ and GenBank formats are on the V3 roadmap.
Is there a sequence length limit?+
General analysis (statistics, ORFs, translation, GC profile) supports up to 500,000 bp. Pairwise alignment is limited to 10,000 bp per sequence for browser performance.
Are ORF predictions biologically validated?+
No. ORF predictions are based purely on computational structure (ATG start, stop codon, minimum 30 aa length). They are candidate regions that require experimental validation.
Which genetic code does the translation use?+
CHOZENLAB uses the Standard Genetic Code (NCBI Genetic Code Table 1) for all translations.
Can I analyze virus or bacterial genomes?+
Yes. CHOZENLAB has been tested on real genomes including HIV-1 reference sequences and bacterial plasmids. Sequences up to 500,000 bp are supported.
Is CHOZENLAB Sequence free to use?+
CHOZENLAB Sequence V2 is free for research and educational use. Commercial licensing enquiries should be directed to contact@chozenlab.com.
How do I cite CHOZENLAB Sequence in a publication?+
Please cite: "CHOZENLAB Sequence V2 — Biomedical Technologies & Life Sciences. Available at: [your deployment URL]. Analysis powered by Biopython (Cock et al., 2009)."
Can I reproduce the results in my own Python environment?+
Absolutely. Every calculation is documented with Biopython code examples in the "Biopython Code Reference" section above. Install Biopython with pip install biopython and run the snippets directly.

Ready to analyze your sequence?

Paste a DNA, RNA, or protein sequence and get a full report in seconds — no login required.

Open Sequence Analyzer →