Search bioRxiv⌕ Search

Biology subjects

Nyasimi, F.

Publications and source records attributed to Nyasimi, F..

2 recordsLinked to original sources

scPrediXcan integrates advances in deep learning and single-cell data into a powerful cell-type-specific transcriptome-wide association study framework

Transcriptome-wide association studies (TWAS) help identify disease causing genes, but often fail to pinpoint disease mechanisms at the cellular level because of the limited sample sizes and sparsity of cell-type-specific expression data. Here we propose scPrediXcan which integrates state-of-the-art deep learning approaches that predict epigenetic features from DNA sequences with the canonical TWAS framework. Our prediction approach, ctPred, predicts cell-type-specific expression with high accuracy and captures complex gene regulatory grammar that linear models overlook. Applied to type 2 diabetes and systemic lupus erythematosus, scPrediXcan outperformed the canonical TWAS framework by identifying more candidate causal genes, explaining more genome-wide association studies (GWAS) loci, and providing insights into the cellular specificity of TWAS hits. Overall, our results demonstrate that scPrediXcan represents a significant advance, promising to deepen our understanding of the cellular mechanisms underlying complex diseases.

genetics↗

On the problem of inflation in transcriptome-wide association studies

Transcription-wide association studies (TWAS) and related methods (xWAS) have been widely adopted in genetic studies to understand molecular traits as mediators between genetic variation and disease. However, the effect of polygenicity on the validity of these mediator-trait association tests has largely been overlooked. Given the widespread polygenicity of complex traits, it is necessary to assess the validity and accuracy of these mediator-trait association tests. We found that for highly polygenic target traits, the standard test based on linear regression is inflated, leading to greatly increased false positives rates, especially in large sample sizes. Here, we show the extent of the inflation as a function of the underlying GWAS sample size and polygenic heritability of the target trait. To address this inflation, we propose an effective variance control method, similar to genomic control, but which allows for a different correction factor for each gene. Using simulated and real data, as well as theoretical derivations, we show that our method yields calibrated false positive rates, outperforming existing approaches. We further demonstrate that methods analogous to TWAS that associate genetic predictors of mediating traits with target traits suffer from similar inflation issues. We advise developers of genetic predictors for molecular traits (including polygenic risk scores, PRS) to compute and provide the necessary inflation parameters to ensure proper false positive control. Finally, we have updated our PrediXcan software package and resources to facilitate this correction for end users.

bioinformatics↗