Search bioRxiv⌕ Search

Biology subjects

Gonzalez-Garcia, L. N.

Publications and source records attributed to Gonzalez-Garcia, L. N..

2 recordsLinked to original sources

New algorithms for accurate and efficient de-novo genome assembly from long DNA sequencing reads

Producing de-novo genome assemblies for complex genomes is possible thanks to long-read DNA sequencing technologies. However, maximizing the quality of assemblies based on long reads is a challenging task that requires the development of specialized data analysis techniques. In this paper, we present new algorithms for assembling long-DNA sequencing reads from haploid and diploid organisms. The assembly algorithm builds an undirected graph with two vertices for each read based on minimizers selected by a hash function derived from the k-mers distribution. Statistics collected during the graph construction are used as features to build layout paths by selecting edges, ranked by a likelihood function that is calculated from the inferred distributions of features on a subset of safe edges. For diploid samples, we integrated a reimplementation of the ReFHap algorithm to perform molecular phasing. The phasing procedure is used to remove edges connecting reads assigned to different haplotypes and to obtain a phased assembly by running the layout algorithm on the filtered graph. We ran the implemented algorithms on PacBio HiFi and Nanopore sequencing data taken from bacteria, yeast, Drosophila, rice, maize, and human samples. Our algorithms showed competitive efficiency and contiguity of assemblies, as well as superior accuracy in some cases, as compared to other currently used software. We expect that this new development will be useful for researchers building genome assemblies for different species.

bioinformatics↗

NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment

Whole-genome alignment allows researchers to understand the genomic structure and variations among the genomes. Approaches based on direct pairwise comparisons of DNA sequences require large computational capacities. As a consequence, pipelines combining tools for orthologous gene identification and synteny have been developed. In this manuscript, we present the latest functionalities implemented in NGSEP 4, to identify orthogroups and perform whole genome alignments. NGSEP implements functionalities for identification of clusters of homologus genes, synteny analysis and whole genome alignment, and visualization. Our results showed that the NGSEP algorithm for ortholog identification has competitive accuracy and better efficiency in comparison to commonly used tools. The implementation also includes a visualization of the whole genome alignment based on synteny of the orthogroups that were identified, and a reconstruction of the pangenome based on frequencies of the orthogroups among the genomes. Finally, our software includes a new graphical user interface. We expect that these new developments will be very useful for several studies in evolutionary biology and population genomics.

bioinformatics↗