Search bioRxiv⌕ Search

Biology subjects

carter, c. W.

Publications and source records attributed to carter, c. W..

3 recordsLinked to original sources

Reduced amino acid substitution matrices find traces of ancient coding alphabets in modern day proteins

All known living systems make proteins from the same twenty canonically-coded amino acids, but this was not always the case. Early genetic coding systems likely operated with a restricted pool of amino acid types and limited means to distinguish between them. Despite this, amino acid substitution models like LG and WAG all assume a constant coding alphabet over time. That makes them especially inappropriate for the aminoacyl-tRNA synthetases (aaRS) - the enzymes that govern translation. To address this limitation, we created a class of substitution models that accounts for evolutionary changes in the coding alphabet size by defining the transition from nineteen states in a past epoch to twenty now. We use a Bayesian phylogenetic framework to improve phylogeny estimation and testing of this two-alphabet hypothesis. The hypothesis was strongly rejected by datasets composed exclusively of "young" eukaryotic proteins. It was generally supported by "old" (aaRS and non-aaRS) proteins whose origins date from before the last universal common ancestor. Standard methods overestimate the divergence ages of proteins that originated under reduced coding alphabets in both simulated and aaRS alignments. The new model reduces this bias substantially. Our findings support the late incorporation of tryptophan into the genetic code (relative to tyrosine) and suggest that isoleucine and valine were once coded interchangeably, forming protein quasispecies. This work provides a robust, seamless framework for reconstructing phylogenies from ancient protein datasets and offers further insights into the dawn of molecular biology.

evolutionary biology↗

Genomic database furnishes a spontaneous example of a functional Class II glycyl-tRNA synthetase urzyme

The chief barrier to studies of how genetic coding emerged is the lack of experimental models for ancestral aminoacyl-tRNA synthetases (AARS). We hypothesized that conserved core catalytic sites could represent such ancestors. That hypothesis enabled engineering functional "urzymes" from TrpRS, LeuRS, and HisRS. We describe here a fourth urzyme, GlyCA, detected in an open reading frame from the genomic record of the arctic fox, Vulpes lagopus. GlyCA is homologous to a bacterial heterotetrameric Class II GlyRS-B. Alphafold2 predicted that the N-terminal 81 amino acids would adopt a 3D structure nearly identical to the HisRS urzyme (HisCA1). We expressed and purified that N-terminal segment. Enzymatic characterization revealed a robust single-turnover burst size and a catalytic rate for ATP consumption well in excess of that previously published for HisCA1. Time-dependent aminoacylation of tRNAGly proceeds at a rate consistent with that observed for amino acid activation. In fact, GlyCA is actually 35 times more active in glycine activation by ATP than the full-length GlyRS-B -subunit dimer. ATP-dependent activation of the 20 canonical amino acids favors Class II amino acids that complement those favored by HisCA and LeuAC. These properties reinforce the notion that urzymes represent the requisite ancestral catalytic activities to implement a reduced genetic coding alphabet.

evolutionary biology↗