Search bioRxivSearch

Biology subjects

Beerenwinkel, N.

Publications and source records attributed to Beerenwinkel, N..

13 recordsLinked to original sources

A stochastic model of metastatic bottleneck predicts patient outcome and therapy response

Metastases are responsible for 90% of cancer-related deaths. Initiation of metastases, where newly seeded tumor cells expand into colonies, presents a tremendous bottleneck to metastasis formation. Despite its clinical importance, our understanding of this process is very limited. Here, we propose a simple stochastic model assuming that the initiating metastatic cells proliferate faster when surrounded by more of their kind. The model quantifies the severity of metastatic bottleneck as the probability that the seeded colony survives. Based on this model, we derive how metastasis occurrence depends on primary tumor size and affects patient outcome. Our predictions agree with epidemiological data for thirteen cancer types. The model predicts that impact of treatment decisions depends both on the primary tumor size and on the severity of the metastatic bottleneck, and that medical interventions that tighten the bottleneck would be much more efficient than therapies that decrease overall tumor burden, such as chemotherapy.

cancer biology

Intra-tumor heterogeneity and clonal exclusivity in renal cell carcinoma

Intra-tumour heterogeneity is the molecular hallmark of renal cancer, and the molecular tumour composition determines the treatment outcome of renal cancer patients. In renal cancer tumourigenesis, in general, different tumour clones evolve over time. We analysed intra-tumour heterogeneity and subclonal mutation patterns in 178 tumour samples obtained from 89 clear cell renal cell carcinoma patients. In an initial discovery phase, whole-exome and transcriptome sequencing data from paired tumour biopsies from 16 ccRCC patients were used to design a gene panel for follow-up analysis. In this second phase, 826 selected genes were targeted at deep coverage in an extended cohort of 89 patients for a detailed analysis of tumour heterogeneity. On average, we found 22 mutations per patient. Pairwise comparison of the two biopsies from the same tumour revealed that on average 62% of the mutations in a patient were detected in one of the two samples. In addition to commonly mutated genes (VHL, PBRM1, SETD2 and BAP1), frequent subclonal mutations with low variant allele frequency (<10%) were observed in TP53 and in mucin coding genes MUC6, MUC16, and MUC3A. Of the 89 ccRCC tumours, 87 (~98%) harboured private mutations, occurring in only one of the paired tumour samples. Clonally exclusive pathway pairs were identified using the WES data set from 16 ccRCC patients. Our findings imply that shared and private mutations significantly contribute to the complexity of differential gene expression and pathway interaction, and might explain clonal evolution of different molecular renal cancer subgroups. Multi-regional sequencing is central for the identification of subclones within ccRCC.

cancer biology

SCIΦ: Single-cell mutation identification via phylogenetic inference

Understanding the evolution of cancer is important for the development of appropriate cancer therapies. The task is challenging because tumors evolve as heterogeneous cell populations with an unknown number of genetically distinct subclones of varying frequencies. Conventional approaches based on bulk sequencing are limited in addressing this challenge as clones cannot be observed directly. Single-cell sequencing holds the promise of resolving the heterogeneity of tumors; however, it has its own challenges including elevated error rates, allelic dropout, and uneven coverage. Here, we develop a new approach to mutation detection in individual tumor cells by leveraging the evolutionary relationship among cells. Our method, called SCI{Phi}, jointly calls mutations in individual cells and estimates the tumor phylogeny among these cells. Employing a Markov Chain Monte Carlo scheme we robustly account for the various sources of noise in single-cell sequencing data. Our approach enables us to reliably call mutations in each single cell even in experiments with high dropout rates and missing data. We show that SCI{Phi} outperforms existing methods on simulated data and applied it to different real-world datasets, namely a whole exome breast cancer as well as a panel acute lymphoblastic leukemia dataset. Availability: https://github.com/cbg-ethz/SCIPhI

bioinformatics

ModulOmics: Integrating Multi-Omics Data to Identify Cancer Driver Modules

The identification of molecular pathways driving cancer progression is a fundamental unsolved problem in tumorigenesis, which can substantially further our understanding of cancer mechanisms and inform the development of targeted therapies. Most current approaches to address this problem use primarily somatic mutations, not fully exploiting additional layers of biological information. Here, we describe ModulOmics, a method to de novo identify cancer driver pathways, or modules, by integrating multiple data types (protein-protein interactions, mutual exclusivity of mutations or copy number alterations, transcriptional co-regulation, and RNA co-expression) into a single probabilistic model. To efficiently search the exponential space of candidate modules, ModulOmics employs a two-step optimization procedure that combines integer linear programming with stochastic search. Across several cancer types, ModulOmics identifies highly functionally connected modules enriched with cancer driver genes, outperforming state-of-the-art methods. For breast cancer subtypes, the inferred modules recapitulate known molecular mechanisms and suggest novel subtype-specific functionalities. These findings are supported by an independent patient cohort, as well as independent proteomic and phosphoproteomic datasets.

cancer biology

Single cell network analysis with a mixture of Nested Effects Models

MotivationNew technologies allow for the elaborate measurement of different traits of single cells. These data promise to elucidate intra-cellular networks in unprecedented detail and further help to improve treatment of diseases like cancer. However, cell populations can be very heterogeneous.\n\nResultsWe developed a mixture of Nested Effects Models (M&NEM) for single-cell data to simultaneously identify different cellular sub-populations and their corresponding causal networks to explain the heterogeneity in a cell population. For inference, we assign each cell to a network with a certain probability and iteratively update the optimal networks and cell probabilities in an Expectation Maximization scheme. We validate our method in the controlled setting of a simulation study and apply it to three data sets of pooled CRISPR screens generated previously by two novel experimental techniques, namely Crop-Seq and Perturb-Seq.\n\nAvailabilityThe mixture Nested Effects Model (M&NEM) is available as the R-package mnem at https://github.com/cbgethz/mnem/.\n\nContactmartin.pirkl@bsse.ethz.ch, niko.beerenwinkel@bsse.ethz.ch\n\nSupplementary informationSupplementary data are available.online.

systems biology

Improved pathway reconstruction from RNA interference screens by exploiting off-target effects

Pathway reconstruction has proven to be an indispensable tool for analyzing the molecular mechanisms of signal transduction underlying cell function. Nested effects models (NEMs) are a class of probabilistic graphical models designed to reconstruct signalling pathways from high-dimensional observations resulting from perturbation experiments, such as RNA interference (RNAi). NEMs assume that the short interfering RNAs (siRNAs) designed to knockdown specific genes are always on-target. However, it has been shown that most siRNAs exhibit strong off-target effects, which further confound the data, resulting in unreliable reconstruction of networks by NEMs. Here, we present an extension of NEMs called probabilistic combinatorial nested effects models (pc-NEMs), which capitalize on the ancillary siRNA off-target effects for network reconstruction from combinatorial gene knockdown data. Our model employs an adaptive simulated annealing search algorithm for simultaneous inference of network structure and error rates inherent to the data. Evaluation of pc-NEMs on simulated data with varying number of phenotypic effects and noise levels demonstrates improved reconstruction compared to classical NEMs. Application to Bartonella henselae infection RNAi screening data yielded an eight node network largely in agreement with previous works, and revealed novel binary interactions of direct impact between established components.\n\nAvailability: The software used for the analysis is freely available as an R package at https://github.com/cbg-ethz/pcNEM.git\n\nContact: niko.beerenwinkel@bsse.ethz.ch

bioinformatics

Integrative inference of subclonal tumour evolution from single-cell and bulk sequencing data

Understanding the evolutionary history and subclonal composition of a tumour represents one of the key challenges in overcoming treatment failure due to resistant cell populations. Most of the current data on tumour genetics stems from short read bulk sequencing data. While this type of data is characterised by low sequencing noise and cost, it consists of aggregate measurements across a large number of cells. It is therefore of limited use for the accurate detection of the distinct cellular populations present in a tumour and the unambiguous inference of their evolutionary relationships. Single-cell DNA sequencing instead provides data of the highest resolution for studying intra-tumour heterogeneity and evolution, but is characterised by higher sequencing costs and elevated noise rates. In this work, we develop the first computational approach that infers trees of tumour evolution from combined single-cell and bulk sequencing data. Using a comprehensive set of simulated data, we show that our approach systematically outperforms existing methods with respect to tree reconstruction accuracy and subclone identification. High fidelity reconstructions are obtained even with a modest number of single cells. We also show that combining single-cell and bulk sequencing data provides more realistic mutation histories for real tumours.

bioinformatics

High-dimensional microbiome interactions shape host fitness

Gut bacteria can affect key aspects of host fitness, such as development, fecundity, and lifespan, while the host in turn shapes the gut microbiome. Microbiomes co-evolve with their hosts and have been implicated in host speciation. However, it is unclear to what extent individual species versus community interactions within the microbiome are linked to host fitness. Here we combinatorially dissect the natural microbiome of Drosophila melanogaster and reveal that interactions between bacteria shape host fitness through life history tradeoffs. We find that the same microbial interactions that shape host fitness also shape microbiome abundances, suggesting a potential evolutionary mechanism by which microbiome communities (rather than just individual species) may be intertwined in co-selection with their hosts. Empirically, we made germ-free flies colonized with each possible combination of the five core species of fly gut bacteria. We measured the resulting bacterial community abundances and fly fitness traits including development, reproduction, and lifespan. The fly gut promoted bacterial diversity, which in turn accelerated development, reproduction, and aging: flies that reproduced more died sooner. From these measurements we calculated the impact of bacterial interactions on fly fitness by adapting the mathematics of genetic epistasis to the microbiome. Host physiology phenotypes were highly dependent on interactions between bacterial species. Higher-order interactions (involving 3, 4, and 5 species) were widely prevalent and impacted both host physiology and the maintenance of gut diversity. The parallel impacts of bacterial interactions on the microbiome and on host fitness suggest that microbiome interactions may be key drivers of evolution.\n\nSignificanceAll animals have associated microbial communities called microbiomes that can influence the physiology and fitness of their host. It is unclear to what extent individual microbial species versus ecology of the microbiome influences fitness of the host. Here we mapped all the possible interactions between individual species of bacteria with each other and with the hosts physiology. Our approach revealed that the same bacterial interactions that shape microbiome abundances also shape host fitness traits. This relationship provides a feedback that may favor the emergence of co-evolving microbiome-host units.

microbiology

Mutational interactions define novel cancer subgroups

Large-scale genomic data can help to uncover the complexity and diversity of the molecular changes that drive cancer progression. Statistical analysis of cancer data from different tissues of origin highlights differences and similarities which can guide drug repositioning as well as the design of targeted and precise treatments. Here, we developed an improved Bayesian network model for tumour mutational profiles and applied it to 8,198 patient samples across 22 cancer types from TCGA. For each cancer type, we identified the interactions between mutated genes, capturing signatures beyond mere mutational frequencies. When comparing mutation networks, we found genes which interact both within and across cancer types. To detach cancer classification from the tissue type we performed de novo clustering of the pancancer mutational profiles based on the Bayesian network models. We found 22 novel clusters which significantly improved survival prediction beyond clinical and histopathological information. The models highlight key gene interactions for each cluster that can be used for genomic stratification in clinical trials and for identifying drug targets within strata.

cancer biology

The geometry of partial fitness orders and an efficient method for detecting genetic interactions

We present an efficient computational approach for detecting genetic interactions from fitness comparison data together with a geometric interpretation using polyhedral cones associated to partial orderings. Genetic interactions are defined by linear forms with integer coefficients in the fitness variables assigned to genotypes. These forms generalize several popular approaches to study interactions, including Fourier-Walsh coefficients, interaction coordinates, and circuits. We assume that fitness measurements come with high uncertainty or are even unavailable, as is the case for many empirical studies, and derive interactions only from comparisons of genotypes with respect to their fitness, i.e. from partial fitness orders. We present a characterization of the class of partial fitness orders that imply interactions, using a graph-theoretic approach. Our characterization then yields an efficient algorithm for testing the condition when certain genetic interactions, such as sign epistasis, are implied. This provides an exponential improvement of the best previously known method. We also present a geometric interpretation of our characterization, which provides the basis for statistical analysis of partial fitness orders and genetic interactions.

evolutionary biology

Context-dependent deposition and regulation of mRNAs in P-bodies

Cells respond to stress by remodeling their transcriptome through transcription and degradation. Xrn1p-dependent degradation in P-bodies is the most prevalent pathway. Yet, P-bodies may facilitate not only decay but also act as storage compartment. However, which and how mRNAs are selected into different degradation pathways and what determines the fate of any given mRNA in P-bodies remain largely unknown. We devised a new method to identify both common and stress-specific mRNA subsets associated with P-bodies. mRNAs targeted for degradation to P-bodies, decayed with different kinetics. Moreover, the localization of a specific set of mRNAs to P-bodies under glucose deprivation was obligatory to prevent decay. Depending on its client mRNA, the RNA binding protein Puf5p either promoted or inhibited decay. The Puf5p-dependent storage of a subset of mRNAs in P-bodies under glucose starvation may be beneficial with respect to chronological lifespan.

cell biology

Inferring Genetic Interactions From Comparative Fitness Data

AO_SCPLOWBSTRACTC_SCPLOWDarwinian fitness is a central concept in evolutionary biology. In practice, however, it is hardly possible to measure fitness for all genotypes in a natural population. Here, we present quantitative tools to make inferences about epistatic gene interactions when the fitness landscape is only incompletely determined due to imprecise measurements or missing observations. We demonstrate that genetic interactions can often be inferred from fitness rank orders, where all genotypes are ordered according to fitness, and even from partial fitness orders. We provide a complete characterization of rank orders that imply higher order epistasis. Our theory applies to all common types of gene interactions and facilitates comprehensive investigations of diverse genetic interactions. We analyzed various genetic systems comprising HIV-1, the malaria-causing parasite Plasmodium vivax, the fungus Aspergillus niger, and the TEM-family of {beta}-lactamase associated with antibiotic resistance. For all systems, our approach revealed higher order interactions among mutations.

evolutionary biology

A statistical test on single-cell data reveals widespread recurrent mutations in tumor evolution

The infinite sites assumption, which states that every genomic position mutates at most once over the lifetime of a tumor, is central to current approaches for reconstructing mutation histories of tumors, but has never been tested explicitly. We developed a rigorous statistical framework to test the assumption with single-cell sequencing data. The framework accounts for the high noise and contamination present in such data. We found strong evidence for recurrent mutations at the same site in 8 out of 9 single-cell sequencing datasets from human tumors. Six cases involved the loss of earlier mutations, five of which occurred at sites unaffected by large scale genomic deletions. Two cases exhibited parallel mutation, including the dataset with the strongest evidence of recurrence. Our results refute the general validity of the infinite sites assumption and indicate that more complex models are needed to adequately quantify intra-tumor heterogeneity.

cancer biology