Search bioRxiv⌕ Search

Biology subjects

Islam, U. I.

Publications and source records attributed to Islam, U. I..

1 recordsLinked to original sources

Scalable PBWT Queries with Minimum-Length SMEMConstraints

Detecting long shared ancestry tracts in large haplotype panels is central to IBD analysis, imputation, and local ancestry inference, and can be approximated computationally by finding Set-Maximal Exact Matches (SMEMs) between sequences. The Positional Burrows-Wheeler Transform (PBWT) provides an efficient index for these panels, yet current methods often enumerate all SMEMs, producing a large number of short, uninformative matches. We introduce Positional Boyer- Moore-Li (PBML), which restricts enumeration to SMEMs occurring in at least k haplotypes and spanning at least L sites (kL-SMEMs). PBML is the first algorithm for computing KL-SMEMs on top of a single compressed run-length encoded PBWT index reusable for any (k, L) without rebuilding. On the 1000 Genomes Project, PBML achieves 4.6x faster query time than {micro}-PBWT and 2.4x over Durbins PBWT with lower memory, scaling to 15.9x over {micro}-PBWT at 16 threads. On a 10,000-haplotype panel from the Tennessee BIG Initiative, a diverse admixed cohort, PBML outperforms {micro}-PBWT by up to 4.7x in k-SMEM finding. By applying both thresholds during traversal, PBML extracts biologically informative, population-shared segments while filtering millions of short matches, a capability not available in current tools. On the BIG panel, in about 10 seconds PBML finds 2,441 long tracts at (k = 50, L = 5000) shared by an average of 60 haplotypes against 1000 queries, significantly reducing the 4.8 million unfiltered SMEMs shared on average by 2 haplotypes. These results establish PBML as a scalable tool for targeted long-range shared ancestry detection in large, diverse panels.

bioinformatics↗