bioRxiv · 10.1101/066381
Codon usage is a stochastic process across genetic codes of the kingdoms of life
Abstract
DNA encodes protein primary structure using 64 different codons to specify 20 different amino acids and a stop signal. To uncover rules of codon use, ranked codon frequencies have previously been analyzed in terms of empirical or statistical relations for a small number of genomes. These descriptions fail on most genomes reported in the Codon Usage Tabulated from GenBank (CUTG) database. Here we model codon usage as a random variable. This stochastic model provides accurate, one-parameter characterizations of 2210 nuclear and mitochondrial genomes represented with > 104 codons/genome in CUTG. We show that ranked codon frequencies are well characterized by a truncated normal (Gaussian) distribution. Most genomes use codons in a nearuniform manner. Lopsided usages are also widely distributed across genomes but less frequent. Our model provides a universal framework for investigating determinants of codon use.
Explore related subjects
Keep this discovery
Bohdan B. Khomtchouk, Claes Wahlestedt, Wolfgang Nonner. 2016-07-27. Codon usage is a stochastic process across genetic codes of the kingdoms of life. https://doi.org/10.1101/066381
Cite the original work for its findings. Save a collection to share your selection of sources.