Search bioRxiv⌕ Search

Biology subjects

Duitama, J.

Publications and source records attributed to Duitama, J..

3 recordsLinked to original sources

NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment

Whole-genome alignment allows researchers to understand the genomic structure and variations among the genomes. Approaches based on direct pairwise comparisons of DNA sequences require large computational capacities. As a consequence, pipelines combining tools for orthologous gene identification and synteny have been developed. In this manuscript, we present the latest functionalities implemented in NGSEP 4, to identify orthogroups and perform whole genome alignments. NGSEP implements functionalities for identification of clusters of homologus genes, synteny analysis and whole genome alignment, and visualization. Our results showed that the NGSEP algorithm for ortholog identification has competitive accuracy and better efficiency in comparison to commonly used tools. The implementation also includes a visualization of the whole genome alignment based on synteny of the orthogroups that were identified, and a reconstruction of the pangenome based on frequencies of the orthogroups among the genomes. Finally, our software includes a new graphical user interface. We expect that these new developments will be very useful for several studies in evolutionary biology and population genomics.

bioinformatics↗

Loss of pod strings in common bean is associated with gene duplication, retrotransposon insertion, and overexpression of PvIND

Regulation of fruit development has been central in the evolution and domestication of flowering plants. In common bean (Phaseolus vulgaris L.), a major global staple crop, the two main economic categories are distinguished by differences in fiber deposition in pods: a) dry beans with fibrous and stringy pods; and b) stringless snap/green beans with reduced fiber deposition, but which frequently revert to the ancestral stringy state. To better understand control of this important trait, we first characterized developmental patterns of gene expression in four phenotypically diverse varieties. Then, using isogenic stringless/revertant pairs of six snap bean varieties, we identified strong overexpression of the common bean ortholog of INDEHISCENT (PvIND) in non-stringy types compared to their string-producing counterparts. Microscopy of these pairs indicates that PvIND overexpression is associated with overspecification of weak dehiscence zone cells throughout the entire pod vascular sheath. No differences in PvIND DNA methylation were correlated with pod string phenotype. Sequencing of a 500 kb region surrounding PvIND in the stringless snap bean cultivar Hystyle revealed that PvIND had been duplicated into two tandem repeats, and that a Ty1-copia retrotransposon was inserted between these tandem repeats, possibly driving PvIND overexpression. Further sequencing of stringless/revertant isogenic pairs and diverse materials indicated that these sequence features had been uniformly lost in revertant types and were strongly predictive of pod phenotype, supporting their role in PvIND overexpression and pod string phenotype. SignificanceFruit dehiscence is a key trait for seed dissemination. In legumes, e.g., common bean, dehiscence relies on the presence of fibers, including pod "strings". Selections during domestication and improvement have decreased (dry beans) or eliminated (snap beans) fibers, but reversion to the fibrous state occurs frequently in snap beans. In this article, we document that fiber loss or gain is controlled by structural changes at the PvIND locus, a homolog of the Arabidopsis INDEHISCENT gene. These changes include a duplication of the locus and insertion/deletion of a retrotransposon, which are associated with significant changes in PvIND expression. Our findings shed light on the molecular basis of unstable mutations and provide potential solutions to an important pod quality issue. Competing Interest StatementThe authors have no competing interests.

plant biology↗

Robust and efficient software for reference-free genomic diversity analysis of GBS data on diploid and polyploid species

Genotype-by-sequencing (GBS) is a widely used cost-effective technique to obtain large numbers of genetic markers from populations. Although a standard reference-based pipeline can be followed to analyze these reads, a reference genome is still not available for a large number of species. Hence, several research groups require reference-free approaches to generate the genetic variability information that can be obtained from a GBS experiment. Unfortunately, tools to perform de-novo analysis of GBS reads are scarce and some of the existing solutions are difficult to operate under different settings generated by the existing GBS protocols. In this manuscript we describe a novel algorithm to perform reference-free variants detection and genotyping from GBS reads. Non-exact searches on a dynamic hash table of consensus sequences allow to perform efficient read clustering and sorting. This algorithm was integrated in the Next Generation Sequencing Experience Platform (NGSEP) to integrate the state-of- the-art variants detector already implemented in this tool. We performed benchmark experiments with three different real populations of plants and animals with different structures and ploidies, and sequenced with different GBS protocols at different read depths. These experiments show that NGSEP has comparable and in some cases better accuracy and always better computational efficiency compared to existing solutions. We expect that this new development will be useful for several research groups conducting population genetic studies in a wide variety of species.

bioinformatics↗