Search bioRxiv⌕ Search

Biology subjects

Kondaramage, D.

Publications and source records attributed to Kondaramage, D..

2 recordsLinked to original sources

PanScan: A Tool for Tertiary Analysis of Human Pangenome Graphs

The genomic representation of populations across the globe is critical to ensuring a comprehensive and equitable human reference. Constructing a pangenome graph reference for different populations is the best approach to addressing local genomic diversities. Although major initiatives across continents are underway to construct pangenome graph references, the field lacks the necessary toolsets for tertiary analysis to characterize telomere-to-telomere (T2T) assemblies and the complexity of haplotypes. PanScan is a bioinformatics software package developed for human pangenome tertiary analysis. It includes multiple modules designed to detect duplicated gene sets from T2T assemblies, identify novel variants and sequences, as well as detect and visualize complex genomic regions through pangenome graph haplotype loops. We have used multiple pangenomes across different populations to assess the tertiary analysis and their accuracy. The tool is designed to streamline tertiary analysis and is compatible with multiple pangenome graph construction algorithms. PanScan is freely available on GitHub (https://github.com/CATG-Github/panscan), where users can provide human pangenome assemblies or VCF files as inputs for automated analyses through command-line operations on Linux systems. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=151 SRC="FIGDIR/small/651685v1_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@16cc212org.highwire.dtl.DTLVardef@13961d0org.highwire.dtl.DTLVardef@44d194org.highwire.dtl.DTLVardef@1b72aa_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

A draft Arab pangenome reference

Pangenomes represent a significant shift from relying on a single reference sequence to a robust set of assemblies, Arab populations remain significantly underrepresented; hence, we present the first Arab Pangenome Reference (APR) utilizing 53 individuals of diverse Arab ethnicities. We assembled nuclear and mitochondrial pangenomes using 35.27X high-fidelity long reads, 54.22X ultralong reads and 65.46X Hi-C reads yielded contiguous haplotype-phased de novo assemblies of exceptional quality, with an average N50 of 124.28 Mb. We discovered 111.96 million base pairs of novel euchromatic sequences absent from existing human pangenomes, the T2T-CHM13, GRCh38 reference human genomes, and other public datasets. We identified 8.94 million population-specific small variants and 235,195 structural variants within the Arab pangenome. We detected 883 gene duplications including 15.06% associated with recessive diseases and 1,436 bp of novel mitochondrial pangenome sequence. Our study provides a valuable resource for future genomic medicine initiatives in Arab population and other global populations.

genomics↗