Search bioRxivSearch

Biology subjects

Lumley, T.

Publications and source records attributed to Lumley, T..

4 recordsLinked to original sources

Chromatin interactions and expression quantitative trait loci reveal genetic drivers of multimorbidities

Clinical studies of non-communicable diseases identify multimorbidities that reflect our relatively limited fixed metabolic capacity. Despite the fact that we have [~]24000 genes, we do not understand the genetic pathways that contribute to the development of multimorbid non-communicable disease. We created a \"multimorbidity atlas\" of traits based on pleiotropy of spatially regulated genes using convex biclustering. Using chromatin interaction and expression Quantitative Trait Loci (eQTL) data, we analysed 20,782 variants (p < 5 x 10-6) associated with 1,351 phenotypes, to identify 16,248 putative eQTL-eGene pairs that are involved in 76,013 short- and long-range regulatory interactions (FDR < 0.05) in different human tissues. Convex biclustering of eGenes that are shared between phenotypes identified complex inter-relationships between nominally different phenotype associated SNPs. Notably, the loci at the centre of these inter-relationships were subject to complex tissue and disease specific regulatory effects. The largest cluster, 40 phenotypes that are related to fat and lipid metabolism, inflammatory disorders, and cancers, is centred on the FADS1-FADS3 locus (chromosome 11). Our novel approach enables the simultaneous elucidation of variant interactions with genes that are drivers of multimorbidity and those that contribute to unique phenotype associated characteristics.

bioinformatics

A Comprehensive Evaluation of the Genetic Architecture of Sudden Cardiac Arrest

BackgroundSudden cardiac arrest (SCA) accounts for 10% of adult mortality in Western populations. While several risk factors are observationally associated with SCA, the genetic architecture of SCA in the general population remains unknown. Furthermore, understanding which risk factors are causal may help target prevention strategies.\n\nMethodsWe carried out a large genome-wide association study (GWAS) for SCA (n=3,939 cases, 25,989 non-cases) to examine common variation genome-wide and in candidate arrhythmia genes. We also exploited Mendelian randomization methods using cross-trait multi-variant genetic risk score associations (GRSA) to assess causal relationships of 18 risk factors with SCA.\n\nResultsNo variants were associated with SCA at genome-wide significance, nor were common variants in candidate arrhythmia genes associated with SCA at nominal significance. Using cross-trait GRSA, we established genetic correlation between SCA and (1) coronary artery disease (CAD) and traditional CAD risk factors (blood pressure, lipids, and diabetes), (2) height and BMI, and (3) electrical instability traits (QT and atrial fibrillation), suggesting etiologic roles for these traits in SCA risk.\n\nConclusionsOur findings show that a comprehensive approach to the genetic architecture of SCA can shed light on the determinants of a complex life-threatening condition with multiple influencing factors in the general population. The results of this genetic analysis, both positive and negative findings, have implications for evaluating the genetic architecture of patients with a family history of SCA, and for efforts to prevent SCA in highrisk populations and the general community.

genetics

A Bayesian approach to multivariate and multilevel modelling with non-random missingness for hierarchical clinical proteomics data

High throughput mass-spectrometry-based proteomics data from clinical studies brings challenges to statistical analysis. The challenges originate from the hierarchical levels of protein abundance data and interactions between clinical study design and experimental design. The non-random missingness of the measurements from a vast amount of information also adds complexity in data analysis. We propose multivariate multilevel models to analyse protein abundances and to handle abundance-dependent missingness within a Bayesian framework. The proposed model enables the variance decomposition at different levels of the data hierarchy and provides shrinkage of protein-level estimates for a group of proteins. A logistic missingness and censored model with informative prior is used to handle incomplete data. Hamiltonian MC/No-U-Turn Sampling and Gibb MCMC algorithms are created to derive the posterior distribution of study parameters; Hamiltonian MC is demonstrated to gain more efficiency for these high-dimensional correlated data. Improvements of the proposed missing data model is compared to the univariate mixed effect model and the multivariate-multilevel model using complete data in a simulated study and a clinical proteomics study. The proposed model framework can be used in other types of data with similar structure and Non Random Missingness mechanism (MNAR).

bioinformatics

Sequence kernel association tests for large sets of markers: tail probabilities for large quadratic forms

The Sequence Kernel Association Test (SKAT) is widely used to test for associations between a phenotype and a set of genetic variants, that are usually rare. Evaluating tail probabilities or quantiles of the null distribution for SKAT requires computing the eigenvalues of a matrix related to the genotype covariance between markers. Extracting the full set of eigenvalues of this matrix (an n x n matrix, for n subjects) has computational complexity proportional to n3. As SKAT is often used when n > 104, this step becomes a major bottleneck in its use in practice. We therefore propose fastSKAT, a new computationally-inexpensive but accurate approximations to the tail probabilities, in which the k largest eigenvalues of a weighted genotype covariance matrix or the largest singular values of a weighted genotype matrix are extracted, and a single term based on the Satterthwaite approximation is used for the remaining eigenval-ues. While the method is not particularly sensitive to the choice of k, we also describe how to choose its value, and show how fastSKAT can automatically alert users to the rare cases where the choice may affect results. As well as providing faster implementation of SKAT, the new method also enables entirely new applications of SKAT, that were not possible before; we give examples grouping variants by topologically assisted domains, and comparing chromosome-wide association by class of histone marker.

bioinformatics