Search bioRxiv⌕ Search

Biology subjects

Zhu, T.-N.

Publications and source records attributed to Zhu, T.-N..

2 recordsLinked to original sources

An Atlas of Linkage Disequilibrium Across Species

Linkage disequilibrium (LD) is a key metric that characterizes populations in flux. To reach a genomic scale LD illustration, which has a substantial computational cost of[O] (nm2), we introduce a framework with two novel algorithms for LD estimation: X-LD, with a time complexity of[O] (n2m) suitable for small sample sizes (n < 104); X-LDR, a stochastic algorithm with a time complexity of[O] (nmB) for biobank-scale data (B iterations); n the sample size, and m the number of SNPs. These methods can refine the entire genome into high-resolution LD grids, such as more than 9 million grids for UK Biobank samples ([~]4.2 million SNPs). The efficient resolution for genome-wide LD leads to intriguing biological discoveries. I) High-resolution LD illustrations revealed how the pericentromeric regions and the HLA region lead to intense and extended LD patterns. II) Two universal LD patterns, identified as Norm I and Norm II patterns, provide insights on the evolutionary history of populations and can also highlight genomic regions of deviation, such as chromosomes 6 and 11 or ncRNA regions. III) The results of our innovative LD decay method aligned with the LD decay scores of 59.5 for Europeans, 60.2 for East Asians, and 33.2 for Africans; correspondingly, the length of the LD was approximately 2.85 Mb, 2.18 Mb, and 1.58 Mb for these three ethnicities. Rare or imputed variants universally increased LD. IV) An unprecedented LD atlas for 25 reference populations contoured interspecies diversity in terms of their Norm I and Norm II LD patterns, highlighting the impact of refined population structure, quality of reference genomes, and uncovered a profound status quo of these populations. The algorithms have been implemented in C++ and are freely available (https://github.com/gc5k/gear2).

genetics↗

Efficient estimation for large-scale linkage disequilibrium patterns of the human genome

In this study, we proposed an efficient algorithm (X-LD) for estimating LD patterns for a genomic grid, which can be of inter-chromosomal scale or of small segments. Compared with conventional methods, the proposed method was significantly faster, dropped from [O] (nm2) to [O] (n2m)--n the sample size and m the number of SNPs, and consequently we were permitted to explore in depth unknown or reveal long-anticipated LD features of the human genome. Having applied the algorithm for 1000 Genome Project (1KG), we found: I) The extended LD, driven by population structure, was universally existed, and the strength of inter-chromosomal LD was about 10% of their respective intra-chromosomal LD in relatively homogeneous cohorts, such as FIN and to nearly 56% in admixed cohort, such as ASW. II) After splitting each chromosome into upmost more than a half million grids, we elucidated the LD of the HLA region was nearly 42 folders higher than chromosome 6 in CEU and 11.58 in ASW; on chromosome 11, we observed that the LD of its centromere was nearly 94.05 folders higher than chromosome 11 in YRI and 42.73 in ASW. III) We uncovered the long-anticipated inversely proportional linear relationship between the length of a chromosome and the strength of chromosomal LD, and their Pearsons correlation was on average over 0.80 for 26 1KG cohorts. However, this linear norm was so far perturbed by chromosome 11 given its more completely sequenced centromere region. Uniquely chromosome 8 of ASW was found most deviated from the linear norm than any other autosomes. The proposed algorithm has been realized in C++ (called X-LD) and available at https://github.com/gc5k/gear2, and can be applied to explore LD features in any sequenced populations.

evolutionary biology↗