Search bioRxivSearch

Biology subjects

Maumus, F.

Publications and source records attributed to Maumus, F..

4 recordsLinked to original sources

Differential retention of transposable element-derived sequences in outcrossing Arabidopsis genomes

Transposable elements (TEs) have initially been viewed as pure genomic parasites but are now recognized as central genome architects and powerful agents of rapid adaptation. A proper evaluation of their evolutionary significance has been hampered by the paucity of short scale phylogenetic comparisons between closely related species. Here, we characterized the dynamics of TE accumulation at the micro-evolutionary scale by comparing two closely related plant species, Arabidopsis lyrata and A. halleri. Joint genome annotation in these two outcrossing species confirmed that both contain two distinct populations of TEs with either recent or old insertion histories. Identification of rare segregating insertions suggests that diverse TE families contribute to the ongoing dynamics of TE accumulation in the two species. TE fragments that have been maintained in both species, i.e. those that are orthologous, tend to be located on average closer to genes than those that are retained in one species only. Moreover, compared to non-orthologous TE insertions, those that are orthologous tend to produce fewer short interfering RNAs, are less heavily methylated when found within or adjacent to genes and these tend to have lower expression levels. These findings suggest that long-term retention of TE insertions reflects their frequent acquisition of adaptive roles and/or the deleterious effects of removing TE insertions when they are close to genes. Overall, our results indicate a rapid evolutionary dynamics of the TE landscape in these two outcrossing species, with an important input of a diverse set of new insertions with variable propensity to resist deletion.

evolutionary biology

Integrative analysis of large scale transcriptome data draws a comprehensive landscape of Phaeodactylum tricornutum functional genome and evolutionary origin of diatoms

Diatoms are one of the most successful and ecologically important groups of eukaryotic phytoplankton in the modern ocean. Deciphering their genomes is a key step towards better understanding of their biological innovations, evolutionary origins, and ecological underpinnings. Here, we have used 90 RNA-Seq datasets from different growth conditions combined with published expressed sequence tags and protein sequences from multiple taxa to explore the genome of the model diatom Phaeodactylum tricornutum, and introduce 1,489 novel genes. The new annotation additionally permitted the discovery for the first time of extensive alternative splicing (AS) in diatoms, including intron retention and exon skipping which increases the diversity of transcripts to regulate gene expression in response to nutrient limitations. In addition, we have used up-to-date reference sequence libraries to dissect the taxonomic origins of diatom genomes. We show that the P. tricornutum genome is replete in lineage-specific genes, with up to 47% of the gene models present only possessing orthologues in other stramenopile groups. Finally, we have performed a comprehensive de novo annotation of repetitive elements showing novel classes of TEs such as SINE, MITE, LINE and TRIM/LARD. This work provides a solid foundation for future studies of diatom gene function, evolution and ecology.

bioinformatics

Tracheophyte genomes keep track of the deep evolution of the Caulimoviridae

Endogenous viral elements (EVEs) are viral sequences that are integrated in the nuclear genomes of their hosts and are signatures of viral infections that may have occurred millions of years ago. The study of EVEs, coined paleovirology, provides important insights into virus evolution. The Caulimoviridae is the most common group of EVEs in plants, although their presence has often been overlooked in plant genome studies. We have refined methods for the identification of caulimovirid EVEs and interrogated the genomes of a broad diversity of plant taxa, from algae to advanced flowering plants. Evidence is provided that almost every vascular plant (tracheophyte), including the most primitive taxa (clubmosses, ferns and gymnosperms) contains caulimovirid EVEs, many of which represent previously unrecognized evolutionary branches. In angiosperms, EVEs from at least one and as many as five different caulimovirid genera were frequently detected, and florendoviruses were the most widely distributed, followed by petuviruses. From the analysis of the distribution of different caulimovirid genera within different plant species, we propose a working evolutionary scenario in which this family of viruses emerged at latest during Devonian era (approx. 320 million years ago) followed by vertical transmission and by several cross-division host swaps.

evolutionary biology

Reconstructing The Gigabase Plant Genome Of Solanum pennellii Using Nanopore Sequencing

Recent updates in sequencing technology have made it possible to obtain Gigabases of sequence data from one single flowcell. Prior to this update, the nanopore sequencing technology was mainly used to analyze and assemble microbial samples1-3. Here, we describe the generation of a comprehensive nanopore sequencing dataset with a median fragment size of 11,979 bp for the wild tomato species Solanum pennellii featuring an estimated genome size of ca 1.0 to 1.1 Gbases. We describe its genome assembly to a contig N50 of 2.5 MB using a pipeline comprising a Canu4 pre-processing and a subsequent assembly using SMARTdenovo. We show that the obtained nanopore based de novo genome reconstruction is structurally highly similar to that of the reference S. pennellii LA7165 genome but has a high error rate caused mostly by deletions in homopolymers. After polishing the assembly with Illumina short read data we obtained an error rate of <0.02 % when assessed versus the same Illumina data. More importantly however we obtained a gene completeness of 96.53% which even slightly surpasses that of the reference S. pennellii genome5. Taken together our data indicate such long read sequencing data can be used to affordably sequence and assemble Gbase sized diploid plant genomes.\n\nRaw data is available at http://www.plabipd.de/portal/solanum-pennellii and has been deposited as PRJEB19787.

genomics