Search bioRxiv⌕ Search

Biology subjects

Maj, C.

Publications and source records attributed to Maj, C..

3 recordsLinked to original sources

Single cell analysis of Barrett's esophagus and carcinoma reveals cell types conferring risk via genetic predisposition

Inherited genetic variants contribute to Barretts esophagus (BE) and esophageal adenocarcinoma (EAC) but it is unknown which cell types are involved in this process. We performed single cell RNA-sequencing of BE, EAC and paired normal tissues and integrated data of a genome-wide association study to determine cell type-specific genetic risk and cellular processes that contribute to BE and EAC. The analysis revealed that EAC development is driven to a greater extent by local cellular processes than BE development. One cell type of BE origin (BE-EAC) and cellular processes that control the differentiation of columnar cells are of particular relevance for EAC development. Further, specific subtypes of fibroblasts and endothelial cells contribute to BE and EAC development, while dendritic cells and CD4+ memory T cells contribute exclusively to BE development. The diagnostic use of markers characterizing the identified cell types and cellular processes should be explored in future for EAC prediction.

cancer biology↗

Boosting polygenic risk scores

Polygenic risk scores (PRS) evaluate the individual genetic liability to a certain trait and are expected to play an increasingly important role in the field of clinical risk stratification. Most often, PRS are estimated based on summary statistics of univariate effects derived from genome-wide association studies. To improve the predictive performance of PRS, it is desirable to fit multivariable models directly on the genetic data. Due to the large and high-dimensional data, a direct application of existing methods is often not feasible and new efficient algorithms are required to overcome the computational burden regarding efficiency and memory demands. We develop an adapted component-wise L2-boosting algorithm to fit genotype data from large cohort studies to continuous outcomes using linear base-learners for the genetic variants. Similar to the snpnet approach implementing lasso regression, the proposed snpboost approach iteratively works on smaller batches of variants. By restricting the set of possible base-learners in each boosting step to variants most correlated with the residuals from previous iterations, the computational efficiency can be substantially increased without losing prediction accuracy. Furthermore, for large-scale data based on various traits from the UK Biobank we show that our method yields competitive prediction accuracy and computational efficiency compared to the snpnet approach. Due to the modular structure of boosting, our framework can be further extended to construct PRS for different outcome data and effect types.

bioinformatics↗

Statistical learning for sparser fine-mapped polygenic models: the prediction of LDL-cholesterol

Polygenic risk scores quantify the individual genetic predisposition regarding a particular trait. We propose and illustrate the application of existing statistical learning methods to derive sparser models for genome-wide data with a polygenic signal. Our approach is based on three consecutive steps. First, potentially informative loci are identified by a marginal screening approach. Then, fine-mapping is independently applied for blocks of variants in linkage disequilibrium, where informative variants are retrieved by using variable selection methods including boosting with probing and stochastic searches with the Adaptive Subspace method. Finally, joint prediction models with the selected variants are derived using statistical boosting. In contrast to alternative approaches relying on univariate summary statistics from genome-wide association studies, our three-step approach enables to select and fit multivariable regression models on large-scale genotype data. Based on UK Biobank data, we develop prediction models for LDL-cholesterol as a continuous trait. Additionally, we consider a recent scalable algorithm for the Lasso. Results show that statistical learning approaches based on fine-mapping of genetic signals result in a competitive prediction performance compared to classical polygenic risk approaches, while yielding sparser risk models that tend to be more robust regarding deviations from the target population.

genetics↗