Search bioRxiv⌕ Search

Biology subjects

Machnik, N.

Publications and source records attributed to Machnik, N..

2 recordsLinked to original sources

Causal inference for multiple risk factors and diseases from genomics data

Statistical causal learning in genome-wide association studies (GWAS) relies on the instrumental variable method of Mendelian Randomization (MR). Currently, an over-whelming number of MR studies purport to show causal relationships among a wide range of risk factors and outcomes. Here, we find that naive application of many recently proposed MR approaches results in numerous null relationships being discovered as highly significant. We show that a well-controlled error rate can be achieved through a graphical inference approach which: (i) selects a set of genetic instrumental variables (IVs) from GWAS summary static controlling for LD, linkage and pleiotropy; (ii) accommodates rare variants and binary outcomes in a principled way; (iii) distinguishes direct from indirect risk factors in very high-dimensional data; and (iv) identifies potential unobserved latent confounding. Only 20 minutes of wall-clock compute time is required for our Causal Inference GWAS (CI-GWAS) approach to jointly analyze a set of 9 common health risk factors, four common complex metabolic disease outcomes and 8.4M genetic variants recorded for 458,747 individuals in the UK Biobank. Genome-wide, we find that very few genetic variants are suitable MR IVs, with only 696 variants remaining when analyzing all traits jointly. While we replicate almost all paths previously found by CAUSE between risk factors and outcomes, we show that only few of these reflect direct adjacencies and that many cannot be distinguished from unmeasured confounding within the UK Biobank data. Our results suggest that well-curated longitudinal records and family data are likely needed to overcome the mixtures of temporal precedence and reverse-causality in biobank data. Our approach provides a first-step toward robust principled screening for potential causal links to understand the underlying nature of phenotypic correlations in biobank data.

genetics↗

SSIM can robustly identify changes in 3D genome conformation maps

We previously presented Comparison of Hi-C Experiments using Structural Similarity (CHESS), an approach that applies the concept of the structural similarity index (SSIM) to Hi-C matrices1, and demonstrated that it could be used to identify both regions with similar 3D chromatin conformation across species, and regions with different chromatin conformation in different conditions. In contrast to the claim of Lee et al.2 that the SSIM output of CHESS is independent of the input data, here we confirm that SSIM depends on both local and global properties of the input Hi-C matrices. We provide two approaches for using CHESS to highlight regions of differential genome organisation for further investigation, and expanded guidelines for choosing appropriate parameters and controls for these analyses.

genomics↗