Search bioRxiv⌕ Search

Biology subjects

Vazquez-Valls, M.

Publications and source records attributed to Vazquez-Valls, M..

3 recordsLinked to original sources

An evaluation of clustering and assembly strategies from Iso-Seq data in the absence of reference genomes in non-model animals

Transcriptome assembly enables the recovery of expressed genes and isoforms, but the optimal strategy for reconstructing transcriptomes from long-read sequencing remains unresolved. In particular, establishing best practices for generating accurate gene models and selecting representative isoforms is essential for comparative genomics, since orthology inference typically requires only the longest isoform per gene model. Here, we systematically compare clustering and de novo assembly methods using PacBio Iso-Seq data from diverse invertebrate lineages with the goal of identifying the most optimal methodology for isoform selection in the absence of dedicated pipelines. We evaluate four approaches: IsoSeq3 (isoseq3 cluster), CD-HIT, RNA-Bloom2 and isONform, all benchmarked against short-read Trinity assemblies. Assembly quality was assessed using BUSCO completeness, short-read mapping rates, coding sequence recovery, longest isoform prediction, and SQANTI3 structural classification. Our results show that CD-HIT clustering at high similarity thresholds ([≥]99%) yields the most complete and coding-rich long-read transcriptomes, rivaling Trinity while avoiding its high redundancy. SQANTI3 classification further confirms that CD-HIT 99 recovers the highest number of full-splice-match transcripts among all methods. Consensus-based methods such as IsoSeq3 and isONform recover fewer single-copy orthologs (mirrored in a lower BUSCO score) and achieve lower mapping rates, while RNA-Bloom2 provides intermediate performance with reduced duplication. Together, these findings establish, to date, CD-HIT as a robust and practical strategy for transcriptome reconstruction from long-read data when genomic references are unavailable. This work provides practical guidance for deriving high-quality gene models and selecting representative isoforms for orthology inference in non-model species.

evolutionary biology↗

Illuminating the functional landscape of the dark proteome across the Animal Tree of Life through natural language processing models

Functional annotation is crucial in biology, but many protein-coding genes remain uncharacterized, especially in non-model organisms. FANTASIA (Functional ANnoTAtion based on embedding space SImilArity) integrates protein language models for large-scale functional annotation. Applied to [~]1,000 animal proteomes, it predicts functions to virtually all proteins, revealing previously uncharacterized functions that enhance our understanding of molecular evolution. FANTASIA is available on GitHub at https://github.com/CBBIO/FANTASIA.

evolutionary biology↗

MATEdb2, a collection of high-quality metazoan proteomes across the Animal Tree of Life to speed up phylogenomic studies

Recent advances in high throughput sequencing have exponentially increased the number of genomic data available for animals (Metazoa) in the last decades, with high-quality chromosome-level genomes being published almost daily. Nevertheless, generating a new genome is not an easy task due to the high cost of genome sequencing, the high complexity of assembly, and the lack of standardized protocols for genome annotation. The lack of consensus in the annotation and publication of genome files hinders research by making researchers lose time in reformatting the files for their purposes but can also reduce the quality of the genetic repertoire for an evolutionary study. Thus, the use of transcriptomes obtained using the same pipeline as a proxy for the genetic content of species remains a valuable resource that is easier to obtain, cheaper, and more comparable than genomes. In a previous study, we presented the Metazoan Assemblies from Transcriptomic Ensembles database (MATEdb), a repository of high-quality transcriptomic and genomic data for the two most diverse animal phyla, Arthropoda and Mollusca. Here, we present the newest version of MATEdb (MATEdb2) that overcomes some of the previous limitations of our database: (1) we include data from all animal phyla where public data is available, (2) we provide gene annotations extracted from the original GFF genome files using the same pipeline. In total, we provide proteomes inferred from high-quality transcriptomic or genomic data for almost 1000 animal species, including the longest isoforms, all isoforms, and functional annotation based on sequence homology and protein language models, as well as the embedding representations of the sequences. We believe this new version of MATEdb will accelerate research on animal phylogenomics while saving thousands of hours of computational work in a plea for open, greener, and collaborative science.

evolutionary biology↗