Search bioRxiv⌕ Search

Biology subjects

Martinez-Redondo, G. I.

Publications and source records attributed to Martinez-Redondo, G. I..

6 recordsLinked to original sources

A punctuated burst of massive genomic rearrangements by chromosome shattering and the origin of non-marine annelids

The genomic basis of cladogenesis and adaptive evolutionary change has intrigued biologists for decades. Here, we show that the tectonics of genome evolution in clitellates, a clade composed of most freshwater and all terrestrial species of the phylum Annelida, is characterised by extensive genome-wide scrambling that resulted in a massive loss of macrosynteny between marine annelids and clitellates. These massive rearrangements included the formation of putative neocentromeres with newly acquired transposable elements and preceded a further period of genome-wide reshaping events, potentially triggered by the loss of genes involved in genome stability and homeostasis of cell division. Notably, while these rearrangements broke short-range interactions observed between Hox genes in marine annelids, they were reformed as long-range interactions in clitellates. Our findings reveal extensive genomic reshaping in clitellates at both the linear (2D) and three-dimensional (3D) levels, suggesting that, unlike in other animal lineages where synteny conservation constrains structural evolution, clitellates exhibit a remarkable tolerance for chromosomal rearrangements. Our study thus suggests that the genomic landscape of Clitellata resulted from a rare burst of genomic changes that ended a long period of stability that persists across large phylogenetic distances.

evolutionary biology↗

Illuminating the functional landscape of the dark proteome across the Animal Tree of Life through natural language processing models

Functional annotation is crucial in biology, but many protein-coding genes remain uncharacterized, especially in non-model organisms. FANTASIA (Functional ANnoTAtion based on embedding space SImilArity) integrates protein language models for large-scale functional annotation. Applied to [~]1,000 animal proteomes, it predicts functions to virtually all proteins, revealing previously uncharacterized functions that enhance our understanding of molecular evolution. FANTASIA is available on GitHub at https://github.com/CBBIO/FANTASIA.

evolutionary biology↗

MATEdb2, a collection of high-quality metazoan proteomes across the Animal Tree of Life to speed up phylogenomic studies

Recent advances in high throughput sequencing have exponentially increased the number of genomic data available for animals (Metazoa) in the last decades, with high-quality chromosome-level genomes being published almost daily. Nevertheless, generating a new genome is not an easy task due to the high cost of genome sequencing, the high complexity of assembly, and the lack of standardized protocols for genome annotation. The lack of consensus in the annotation and publication of genome files hinders research by making researchers lose time in reformatting the files for their purposes but can also reduce the quality of the genetic repertoire for an evolutionary study. Thus, the use of transcriptomes obtained using the same pipeline as a proxy for the genetic content of species remains a valuable resource that is easier to obtain, cheaper, and more comparable than genomes. In a previous study, we presented the Metazoan Assemblies from Transcriptomic Ensembles database (MATEdb), a repository of high-quality transcriptomic and genomic data for the two most diverse animal phyla, Arthropoda and Mollusca. Here, we present the newest version of MATEdb (MATEdb2) that overcomes some of the previous limitations of our database: (1) we include data from all animal phyla where public data is available, (2) we provide gene annotations extracted from the original GFF genome files using the same pipeline. In total, we provide proteomes inferred from high-quality transcriptomic or genomic data for almost 1000 animal species, including the longest isoforms, all isoforms, and functional annotation based on sequence homology and protein language models, as well as the embedding representations of the sequences. We believe this new version of MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open, greener, and collaborative science.

evolutionary biology↗

Decoding proteome functional information in model organisms using protein language models.

Protein language models have been tested and proved to be reliable when used on curated datasets but have not yet been applied to full proteomes. Accordingly, we tested how two different machine learning based methods performed when decoding functional information from the proteomes of selected model organisms. We found that protein Language Models are more precise and informative than Deep Learning methods for all the species tested and across the three gene ontologies studied, and that they better recover functional information from transcriptomics experiments. The results obtained indicate that these Language Models are likely to be suitable for large scale annotation and downstream analyses, and we recommend a guide for their use.

bioinformatics↗

Parallel duplication and loss of aquaporin-coding genes during the 'out of the sea' transition paved the way for animal terrestrialization

One of the most important physiological challenges animals had to overcome during terrestrialization (i.e., the transition from sea to land) is water loss, which alters their osmotic and hydric homeostasis. Aquaporins are a superfamily of membrane water transporters heavily involved in osmoregulatory processes. Their diversity and evolutionary dynamics in most animal lineages remain unknown, hampering our understanding of their role in marine-terrestrial transitions. Here, we interrogated aquaporin gene repertoire evolution across the main terrestrial animal lineages. We annotated aquaporin-coding genes in genomic data from 458 species from 7 animal phyla where terrestrialization episodes occurred. We then explored aquaporin gene evolutionary dynamics to assess differences between terrestrial and aquatic species through phylogenomics and phylogenetic comparative methods. Our results revealed parallel aquaporin-coding gene duplications in aquaporins during the transition from marine to non-marine environments (e.g., brackish, freshwater and terrestrial), rather than from aquatic to terrestrial ones, with some notable duplications in ancient lineages. Contrarily, we also recovered a significantly lower number of superaquaporin genes in terrestrial arthropods, suggesting that more efficient oxygen homeostasis in land arthropods might be linked to a reduction in this type of aquaporins. Our results thus indicate that aquaporin-coding gene duplication and loss might have been one of the key steps towards the evolution of osmoregulation across animals, facilitating the out of the sea transition and ultimately the colonisation of land.

evolutionary biology↗

MATEdb, a data repository of high-quality metazoan transcriptome assemblies to accelerate phylogenomic studies

AO_SCPLOWBSTRACTC_SCPLOWWith the advent of high throughput sequencing, the amount of genomic data available for animals (Metazoa) species has bloomed over the last decade, especially from transcriptomes due to lower sequencing costs and easier assembling process compared to genomes. Transcriptomic data sets have proven useful for phylogenomic studies, such as inference of phylogenetic interrelationships (e.g., species tree reconstruction) and comparative genomics analyses (e.g., gene repertoire evolutionary dynamics). However, these data sets are often analyzed following different analytical pipelines, particularly including different software versions, leading to potential methodological biases when analyzed jointly in a comparative framework. Moreover, these analyses are computationally expensive and not affordable for a large part of the scientific community. More importantly, assembled transcriptomes are usually not deposited in public databases. Furthermore, the quality of these data sets is hardly ever taken into consideration, potentially impacting subsequent analyses such as orthology and phylogenetic or gene repertoire evolution inference. To alleviate these issues, we present Metazoan Assemblies from Transcriptomic Ensembles (MATEdb), a curated database of 335 high-quality transcriptome assemblies from different animal phyla analyzed following the same pipeline. The repository is composed, for each species, of (1) a de novo transcriptome assembly, (2) its candidate coding regions within transcripts (both at the level of nucleotide and amino acid sequences), (3) the coding regions filtered using their contamination profile (i.e., only metazoan content), (4) the longest isoform of the amino acid candidate coding regions, (5) the gene content completeness score as assessed against the BUSCO database, and (6) an orthology-based gene annotation. We complement the repository with gene annotations from high-quality genomes, which are often not straightforward to obtain from individual sequencing projects, totalling 423 high-quality genomic and transcriptomic data sets. We invite the community to provide suggestions for new data sets and new annotation features to be included in subsequent versions, that will be analyzed following the same pipeline and be permanently stored in public repositories. We believe that MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open and collaborative science.

evolutionary biology↗