Search bioRxiv⌕ Search

Biology subjects

Alkhnbashi, O.

Publications and source records attributed to Alkhnbashi, O..

4 recordsLinked to original sources

VDB: The Arab Variation and Disease Burden Database

MotivationPublic genomic databases are crucial to precision medicine but often lack representation from Arab populations, which have distinct genetic structures due to high consanguinity rates. ResultsWe present the Arab Variation and Disease Burden Database (AVDB), derived from 1,194 exomes of Emirati individuals, comprising 2,481 curated variants across 850 genes. AVDB provides pathogenicity classification, carrier frequencies, and gene-level at-risk couple rates (ARCR), reaching 21% under first-cousin mating models. It integrates machine learning for gene prioritization and offers a customizable panel builder. Compared to global panels, AVDB captures more regional burden, filling a major gap in population-specific screening. The resource supports genomic diagnostics and health policy for Arab populations and is freely accessible at https://avdb-arabgenome.ae. Contactomer.alkhnbashi@dubaihealth.ae and ahmad.tayoun@dubaihealth.ae Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

An evolutionary approach to predict the orientation of CRISPR arrays

CRISPR-Cas is a defense system of bacteria and archaea against phages. Parts of the foreign DNA, called spacers, are incorporated into the CRISPR array which constitutes the immune memory. The orientation of CRISPR arrays is crucial for analyzing and understanding the functionality of CRISPR systems and their targets. Several methods have been developed to identify the orientation of a CRISPR array. To predict the orientation, different methods use different features such as the repeat sequences between the spacers, the location of the leader sequence, the Cas genes, or PAMs. However, those features are often not sufficient to predict the orientation with certainty, or different methods disagree. Remarkably, almost all CRISPR systems have been found to insert spacers in a polarized manner at the leader end of the array. We introduce CRISPR-evOr, a method that leverages the resulting patterns to predict the acquisition orientation for (a group of) CRISPR arrays by reconstructing and comparing the likelihood of their evolutionary history with respect to both possible acquisition orientations. The new method is independent of Cas type, leader existence and location, and transcription orientation. CRISPR-evOr is thus particularly useful for arrays that other CRISPR orientation tools cannot predict confidently and to verify or resolve conflicting predictions from existing tools. CRISPR-evOr currently confidently predicts the orientation of 28.3% of the arrays in the considered subset of CRISPRCasdb, which other tools like CRISPRDirection and CRISPRstrand cannot reliably orient. As our tool leverages evolutionary information we expect this percentage to grow in the future when more closely related arrays will be available. Additionally, CRISPR-evOr provides confident decisions for rare subtypes of CRISPR arrays, where knowledge about repeats and leaders and their orientation is limited. Author SummarySome bacteria and archaea possess a CRISPR-Cas defense system, which protects them against phages and mobile genetic elements. This system adapts to new threats by incorporating small fragments of their DNA, so called spacers, in a CRISPR array. Remarkably, the acquisition of new spacers is polarized at one end of this array. To understand how this immune system functions, it is essential to know the orientation of these arrays. Many existing tools try to determine orientation using genetic markers, but these methods are often unreliable or disagree with one another. In our work, we developed a new method that predicts the end at which new spacers are inserted, by considering the evolutionary history of a group of related CRISPR arrays. Unlike other tools, our approach is less reliant on specific genetic markers and can be applied broadly across many types of CRISPR systems. We show that it can confidently determine the orientation of a large number of arrays that other methods cannot resolve. This provides a new way to predict the array orientation, which is particularly useful for rare CRISPR types. Our evolutionary approach will become even more powerful as more genetic data becomes available.

bioinformatics↗

PanScan: A Tool for Tertiary Analysis of Human Pangenome Graphs

The genomic representation of populations across the globe is critical to ensuring a comprehensive and equitable human reference. Constructing a pangenome graph reference for different populations is the best approach to addressing local genomic diversities. Although major initiatives across continents are underway to construct pangenome graph references, the field lacks the necessary toolsets for tertiary analysis to characterize telomere-to-telomere (T2T) assemblies and the complexity of haplotypes. PanScan is a bioinformatics software package developed for human pangenome tertiary analysis. It includes multiple modules designed to detect duplicated gene sets from T2T assemblies, identify novel variants and sequences, as well as detect and visualize complex genomic regions through pangenome graph haplotype loops. We have used multiple pangenomes across different populations to assess the tertiary analysis and their accuracy. The tool is designed to streamline tertiary analysis and is compatible with multiple pangenome graph construction algorithms. PanScan is freely available on GitHub (https://github.com/CATG-Github/panscan), where users can provide human pangenome assemblies or VCF files as inputs for automated analyses through command-line operations on Linux systems. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=151 SRC="FIGDIR/small/651685v1_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@16cc212org.highwire.dtl.DTLVardef@13961d0org.highwire.dtl.DTLVardef@44d194org.highwire.dtl.DTLVardef@1b72aa_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Comprehensive Analysis of CRISPR Array Repeat Mutations Reveals Subtype-Specific Patterns and Links to Spacer Dynamics

CRISPR-Cas systems are adaptive immune mechanisms in bacteria and archaea that protect against invading genetic elements by integrating short fragments of foreign DNA into CRISPR arrays. These arrays consist of repetitive sequences interspersed with unique spacers, guiding Cas proteins to recognize and degrade matching nucleic acids. The integrity of these repeat sequences is crucial for the proper function of CRISPR-Cas systems, yet their mutational dynamics remain poorly understood. In this study, we analyzed 56,343 CRISPR arrays across 25,628 diverse prokaryotic genomes to assess the mutation patterns in CRISPR array repeat sequences within and across different CRISPR subtypes. Our findings reveal, as expected to some extent, that mutation frequency is substantially higher in terminal repeat sequences compared to internal repeats consistently across system types. However, the mutation patterns exhibit an unexpected amount of variation among different CRISPR subtypes, suggesting that selective pressures and functional constraints shape repeat sequence evolution in distinct ways. Understanding these mutation dynamics provides insights into the stability and adaptability of CRISPR arrays across diverse bacterial and archaeal lineages. Additionally, we elucidate a novel relationship between repeat mutations and spacer dynamics, demonstrating that hotspots for terminal repeat mutations coincide with regions exhibiting spacer conservation. This observation corroborates recent findings by Fehrenbach et al. (2024) indicating that spacer deletions occur at a frequency 374 times greater than that of mutations and are significantly influenced by repeat misalignment. Our findings suggest that repeat mutations play a pivotal role in spacer retention or loss, or vice versa, thereby highlighting an evolutionary trade-off between the stability and adaptability of CRISPR arrays.

bioinformatics↗