Search bioRxiv⌕ Search

Biology subjects

Beentjes, S. V.

Publications and source records attributed to Beentjes, S. V..

3 recordsLinked to original sources

Systematic identification of context-dependent gene essentiality in Glioblastoma: The GBM-CoDE platform

Glioblastoma (GBM) is a heterogeneous and aggressive brain tumour that is invariably fatal despite maximal treatment. Genetic or transcriptomic biomarkers could be used to stratify patients for treatments, however, pairing biomarkers with appropriate therapeutic targets is challenging. Consequently, therapeutics have not yet been optimised for specific GBM patient subsets. Here we integrate genome-wide CRISPR/Cas9 knockout screening and genetic-annotation data for 60 distinct patient-derived, IDHwildtype, adult GBM cell lines, quantifying the essentiality of 15,145 genes. We describe a novel method using Targeted Learning, to estimate the effect size of GBM-relevant biomarkers on context-dependent gene essentiality (GBM-CoDE). We derive multiple target-biomarker pair hypotheses, which we release in an accessible platform to accelerate translation to biomarker-stratified clinical trials. Two of these (WWTR1 with EGFR mutation/amplification, and VRK1 with VRK2 expression suppression) have been validated in GBM, implying that our additional novel findings may be valid. Our method is readily translatable to other cancers of unmet need.

bioinformatics↗

High order expression dependencies finely resolve cryptic states and subtypes in single cell data

AO_SCPLOWBSTRACTC_SCPLOWSingle cells are typically typed by clustering in reduced dimensional transcriptome space. Here we introduce Stator, a novel method, workflow and app that reveals cell types, subtypes and states without relying on local proximity of cells in gene expression space. Rather, Stator derives higher-order gene expression dependencies from a sparse gene-by-cell expression matrix. From these dependencies the method multiply labels the same single cell according to type, sub-type and state (activation, differentiation or cell cycle sub-phase). By applying the method to data from mouse embryonic brain, and human healthy or diseased liver, we show how Stator first recapitulates other methods cell type labels, and then reveals combinatorial gene expression markers of cell type, state, and disease at higher resolution. By allowing multiple state labels for single cells we reveal cell type fates of embryonic progenitor cells and liver cancer states associated with patient survival.

molecular biology↗

Dispensing with unnecessary assumptions in population genetics analysis

Parametric assumptions in population genetics analysis - including linearity, sources of population stratification and additivity of variance as part of a Gaussian noise - are often made, yet their (approximate) validity depends on variant and traits of interest, as well as genetic ancestry and population dependence structure of the sample cohort. We present a unified statistical workflow, called TarGene, for targeted estimation of effect sizes, as well as two-point and higher-order epistatic interactions of genomic variants on polygenic traits, which dispenses with these unnecessary assumptions. Our approach is founded on Targeted Learning, a framework for estimation that integrates mathematical statistics, machine learning and causal inference. TarGene maximises power whilst simultaneously maximising control over false discoveries by: (i) guaranteeing optimal bias-variance trade-off, (ii) taking into account potential covariate non-linearities, sources of population stratification and dependence structure, and (iii) detecting genetic non-linearities. The necessity of this model-independent approach is demonstrated via extensive simulations. We validate the effectiveness of our method by reproducing previously verified effect sizes on UK Biobank data, whilst simultaneously discovering non-linear effect sizes of additional allelic copies on trait or disease, in a PheWAS study involving 781 traits. Specifically, we demonstrate genetic non-linearity at the FTO locus is significant for 54 traits in this study. We further find three pairs of epistatic loci associated with skin color that have been previously reported to be associated with hair color. Finally, we illustrate how TarGene can be used to investigate higher-order interactions using three variants linked to the vitamin D receptor complex. TarGene provides a platform for comparative analyses across biobanks, or integration of multiple biobanks and heterogeneous populations to simultaneously increase power and control for type I errors, whilst taking into account population stratification and complex dependence structures.

genetics↗