Search bioRxiv⌕ Search

Biology subjects

Mattick, J. S. A.

Publications and source records attributed to Mattick, J. S. A..

2 recordsLinked to original sources

Group II Introns in Archaeal Genomes and the Evolutionary Origin of Eukaryotic Spliceosomal Introns

A key attribute of eukaryotic genomes is the presence of abundant spliceosomal introns that break up many protein-coding genes into multiple exons and must be spliced out during the process of gene expression. These introns are believed to be evolutionarily derived from group II introns, which are known to be widespread in bacteria. One prominent hypothesis is that the spliceosomal intron arose after the endosymbiotic origin of the mitochondrion, as a consequence of transfer of genes containing group II introns from the organelle to nuclear genome; in this model, transfer of group II introns into the ancestral eukaryotic genome set the stage for evolution of the spliceosomal form. However, the recent discovery and sequencing of asgard archaea -- the closest archaeal relatives of extant eukaryotes -- has shed significant light on the composition of the early eukaryotic genome and calls that model into question. Using sequence analysis and structural modeling, we show here the presence of group II intron maturases in the genomes of Heimdallarchaeia and other asgard archaea, and demonstrate by phylogenetic inference that these are closely related to both eukaryotic mitochondrial group II intron maturases and the spliceosome protein PRP8. This suggests that the first intron-containing eukaryotic common ancestor (FIECA) inherited selfish group II introns from its ancestral archaeal genome - the progenitor of the nuclear genome - rather than from the mitochondrial endosymbiont. These observations suggest that the spread and diversification of introns may have occurred independently of the acquisition of the mitochondrion. To better understand the context for intron evolution, we investigate the broader occurrence of group II introns in archaea, identify archaeal clades enriched in group II introns, and perform structural modeling to examine the relationship between the archaeal group II intron maturase and the eukaryotic spliceosome. We propose a model of intron acquisition and expansion during early eukaryotic evolution that places the spread of introns prior to the acquisition of mitochondria, possibly facilitated by the separation of transcription and translation afforded by the nucleus.

evolutionary biology↗

Deciphering Bacterial and Archaeal Transcriptional Dark Matter and Its Architectural Complexity

Transcripts are potential therapeutic targets, yet bacterial transcripts remain biological dark matter with uncharacterized biodiversity. We developed and applied an algorithm to predict transcripts for Escherichia coli K12 and E2348/69 strains (Bacteria:gamma-Proteobacteria) with newly generated ONT direct RNA sequencing data while predicting transcripts for Listeria monocytogenes strains Scott A and RO15 (Bacteria:Firmicute), Pseudomonas aeruginosa strains SG17M and NN2 strains (Bacteria:gamma-Proteobacteria), and Haloferax volcanii (Archaea:Halobacteria) using publicly available data. From >5 million E. coli K12 ONT direct RNA sequencing reads, 2,484 mRNAs are predicted and contain more than half of the predicted E. coli proteins. While the number of predicted transcripts varied by strain based on the amount of sequence data used for the predictions, across all strains examined, the average size of the predicted mRNAs is 1.6-1.7 kbp while the median size of the predicted bacterial 5-and 3-UTRs are 30-90 bp. Given the lack of bacterial and archaeal transcript annotation, most predictions are of novel transcripts, but we also predicted many previously characterized mRNAs and ncRNAs, including post-transcriptionally generated transcripts and small RNAs associated with pathogenesis in the E. coli E2348/69 LEE pathogenicity islands. We predicted small transcripts in the 100-200 bp range as well as >10 kbp transcripts for all strains, with the longest transcript for two of the seven strains being the nuo operon transcript, and for another two strains it was a phage/prophage transcript. This quick, easy, inexpensive, and reproducible method will facilitate the presentation of operons, transcripts, and UTR predictions alongside CDS and protein predictions in bacterial genome annotation as important resources for the research community.

genomics↗