DNA coding capacity calculator
A DNA open reading frame encodes one amino acid per three base pairs, with one codon spent on the stop: a 1,000 bp ORF encodes 332 residues, or roughly a 36.5 kDa protein at the 110 Da-per-residue average. Reversed, a 50 kDa protein (~455 residues) needs about 1,368 bp of coding sequence.
Result
Three bp per codon; one codon is the stop; ~110 Da per amino-acid residue (Promega BioMath convention). A 1,000 bp ORF encodes 332 residues ≈ 36.5 kDa.
Doing this for a real construct?
Talindrew runs it on your actual sequence — editor, cloning wizards and analysis in the browser, free.
Formula
codons = ⌊bp / 3⌋ · residues = codons − 1 (stop) · protein kDa ≈ residues × 110 / 1000
Source: Promega BioMath coding-capacity convention (110 Da average residue)
Worked example
Given: A 1,000 bp open reading frame
- 1.Codons: 1,000 / 3 = 333 (rounded down).
- 2.One codon is the stop → 332 amino acids.
- 3.Protein mass: 332 × 110 Da = 36.5 kDa.
1,000 bp encodes ~332 aa ≈ 36.5 kDa of protein
How the calculation flows
Units & constants
| Codon size | 3 bp; no introns assumed (prokaryotic/cDNA arithmetic) |
|---|---|
| Average residue mass | 110 Da (range ~57 Gly – 186 Trp) |
| Rule of thumb | 1 kb ≈ 37 kDa of protein; 1 kDa protein ≈ 27 bp |
| Start codon | Counted as residue 1 (Met, usually retained or cleaved in vivo) |
| Price | Free |
How much DNA does a 50 kDa protein need?
About 1.4 kb. 50 kDa at 110 Da per residue is ~455 amino acids; add the stop codon and multiply by three: 456 × 3 = 1,368 bp. The estimate ignores tags and linkers — a His₆ tag adds 18 bp plus its linker, and fusion partners add their own coding length.
When does the estimate break down?
The arithmetic is exact for the codon count but the mass is an average:
- Unusual composition (Gly/Ala-rich or Trp-rich proteins) shifts real MW by several percent — compute the exact mass from the sequence when it matters.
- Eukaryotic genes with introns encode far less protein per genomic bp; use the mRNA/cDNA length.
- Post-translational modifications (glycosylation especially) add mass invisible to this calculation.
Frequently asked questions
- Why subtract one codon?
- The stop codon terminates translation without adding a residue, so an ORF of N codons yields N − 1 amino acids. For long proteins the difference is negligible; for short peptides it is not.
- Is 110 Da per residue accurate?
- It is the accepted average across typical protein composition (Promega's convention; some sources use 110–115). Individual residues range from 57 (Gly) to 186 (Trp), so exact work should compute mass from the actual sequence — the in-app protein tools do this.
Related calculators & guides
Use this calculator inside the app.
Free to use. No download required. Sign in and you land right back on this calculator.