Search bioRxivSearch

Biology subjects

Gozashti, L.

Publications and source records attributed to Gozashti, L..

3 recordsLinked to original sources

Massive intron gain in the most intron-rich eukaryotes is driven by introner-like transposable elements of unprecedented diversity and flexibility

Spliceosomal introns, which interrupt nuclear genes and are removed from RNA transcripts by machinery termed spliceosomes, are ubiquitous features of eukaryotic nuclear genes [1]. Patterns of spliceosomal intron evolution are complex, with some lineages exhibiting virtually no intron creation while others experience thousands of intron gains [2-5]. One possibility is that this punctate phylogenetic distribution is explained by intron creation by Introner-Like Elements (ILEs), transposable elements capable of creating introns, with only those lineages harboring ILEs undergoing massive intron gain [6-10]. However, ILEs have been reported in only four lineages. Here we study intron evolution in dinoflagellates. The remarkable fragmentation of nuclear genes by spliceosomal introns reaches its apex in dinoflagellates, which have some twenty introns per gene [11,12]. Despite this, almost nothing is known about the molecular and evolutionary mechanisms governing dinoflagellate intron evolution. We reconstructed intron evolution in five dinoflagellate genomes, revealing a dynamic history of intron loss and gain. ILEs are found in 4/5 studied species. In one species, Polarella glacialis, we find an unprecedented diversity of ILEs, with ILE insertion leading to creation of some 12,253 introns, and with 15 separate families of ILEs accounting for at least 100 introns each. These ILE families range in mobilization mechanism, mechanism of intron creation, and flexibility of mechanism of intron creation. Comparison within and between ILE families provides evidence that biases in so-called intron phase, the distribution of introns relative to codon periodicity, are driven by ILE insertion site requirements [9,13,14]. Finally, we find evidence for multiple additional transformations of the spliceosomal system in dinoflagellates, including widespread loss of ancestral introns, and alterations in required, tolerated and favored splice motifs. These results reveal unappreciated intron creating elements diversity and spliceosomal evolutionary capacity, and suggest complex evolutionary dependencies shaping genome structures.

genomics

Ultrafast Sample Placement on Existing Trees (UShER) Empowers Real-Time Phylogenetics for the SARS-CoV-2 Pandemic

As the SARS-CoV-2 virus spreads through human populations, the unprecedented accumulation of viral genome sequences is ushering a new era of "genomic contact tracing" - that is, using viral genome sequences to trace local transmission dynamics. However, because the viral phylogeny is already so large - and will undoubtedly grow many fold - placing new sequences onto the tree has emerged as a barrier to real-time genomic contact tracing. Here, we resolve this challenge by building an efficient, tree-based data structure encoding the inferred evolutionary history of the virus. We demonstrate that our approach improves the speed of phylogenetic placement of new samples and data visualization by orders of magnitude, making it possible to complete the placements under real-time constraints. Our method also provides the key ingredient for maintaining a fully-updated reference phylogeny. We make these tools available to the research community through the UCSC SARS-CoV-2 Genome Browser to enable rapid cross-referencing of information in new virus sequences with an ever-expanding array of molecular and structural biology data. The methods described here will empower research and genomic contact tracing for laboratories worldwide. Software AvailabilityUSHER is available to users through the UCSC Genome Browser at https://genome.ucsc.edu/cgi-bin/hgPhyloPlace. The source code and detailed instructions on how to compile and run UShER are available from https://github.com/yatisht/usher.

genomics

Stability of SARS-CoV-2 Phylogenies

The SARS-CoV-2 pandemic has led to unprecedented, nearly real-time genetic tracing due to the rapid community sequencing response. Researchers immediately leveraged these data to infer the evolutionary relationships among viral samples and to study key biological questions, including whether host viral genome editing and recombination are features of SARS-CoV-2 evolution. This global sequencing effort is inherently decentralized and must rely on data collected by many labs using a wide variety of molecular and bioinformatic techniques. There is thus a strong possibility that systematic errors associated with lab-specific practices affect some sequences in the repositories. We find that some recurrent mutations in reported SARS-CoV-2 genome sequences have been observed predominantly or exclusively by single labs, co-localize with commonly used primer binding sites and are more likely to affect the protein coding sequences than other similarly recurrent mutations. We show that their inclusion can affect phylogenetic inference on scales relevant to local lineage tracing, and make it appear as though there has been an excess of recurrent mutation and/or recombination among viral lineages. We suggest how samples can be screened and problematic mutations removed. We also develop tools for comparing and visualizing differences among phylogenies and we show that consistent clade- and tree-based comparisons can be made between phylogenies produced by different groups. These will facilitate evolutionary inferences and comparisons among phylogenies produced for a wide array of purposes. Building on the SARS-CoV-2 Genome Browser at UCSC, we present a toolkit to compare, analyze and combine SARS-CoV-2 phylogenies, find and remove potential sequencing errors and establish a widely shared, stable clade structure for a more accurate scientific inference and discourse. ForewordWe wish to thank all groups that responded rapidly by producing these invaluable and essential sequence data. Their contributions have enabled an unprecedented, lightning-fast process of scientific discovery---truly an incredible benefit for humanity and for the scientific community. We emphasize that most lab groups with whom we associate specific suspicious alleles are also those who have produced the most sequence data at a time when it was urgently needed. We commend their efforts. We have already contacted each group and many have updated their sequences. Our goal with this work is not to highlight potential errors, but to understand the impacts of these and other kinds of highly recurrent mutations so as to identify commonalities among the suspicious examples that can improve sequence quality and analysis going forward.

genomics