Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Discrete mutations in colorectal cancer correlate with defined microbial communities in the tumor microenvironment

Variation in the gut microbiome has been linked to colorectal cancer (CRC), as well as to host genetic variation. However, we do not know whether, in addition to baseline host genetics, somatic mutational profiles in CRC tumors interact with the surrounding tumor microbiome, and if so, whether these changes can be used to understand microbe-host interactions with potential functional biological relevance. Here, we characterized the association between CRC microbial communities and tumor mutations using microbiome profiling and whole-exome sequencing in 44 pairs of tumors and matched normal tissues. We found statistically significant associations between loss-of-function mutations in tumor genes and shifts in the abundances of specific sets of bacterial taxa, suggestive of potential functional interaction. This correlation allows us to statistically predict interactions between loss-of-function tumor mutations in cancer-related genes and pathways, including MAPK and Wnt signaling, solely based on the composition of the microbiome. These results can serve as a starting point for fine-grained exploration of the functional interactions between discrete alterations in tumor DNA and proximal microbial communities in CRC. In addition, these findings can lead to the development of improved microbiome-based CRC screening methods, as well as individualized microbiota-targeting therapies.\n\nAuthor summaryAlthough the gut microbiome - the collection of microorganisms that inhabit our gastrointestinal tract - has been implicated in colorectal cancer, colorectal tumors are caused by genetic mutations in host DNA. Here, we explored whether various mutations in colorectal tumors are correlated with specific changes in the bacterial communities that live in and on these tumors. We find that the genes and biological pathways that are mutated in tumors are correlated with variation in the composition of the microbiome. In fact, these changes in the microbiome are consistent enough that we can use them to statistically predict tumor mutations solely based on the microbiome. Our results may be used to understand the roles of specific microbes in CRC biology, and could also be the starting point of microbiome-based diagnostics for not only detection of CRC, but characterization of tumor mutational profiles.

cancer biology

Leveraging collective regulatory effects of long-range DNA methylations to predict gene expressions and estimate their effects on phenotypes in cancer

DNA methylation of various genomic regions plays an important role in regulating gene expression in diverse biological contexts. However, most genome-wide studies have focused on the effect of 1) methylation in cis, not in trans and 2) a single CpG, not the collective effects of multiple CpGs, on gene expression. In this study, we developed a statistical machine learning model, geneEXPLORER (gene expression prediction by long-range epigenetic regulation), that quantifies the collective effects of both cis- and trans- methylations on gene expression. By applying geneEXPLORER to The Cancer Genome Atlas (TCGA) breast and lung cancer data, we found that most genes are affected by methylations of as much as 10Mb from promoter regions or more, and the long-range methylation explains 50% of the variation in gene expression on average, far greater than cis-methylation. The highly predictive genes are related to breast cancer, especially oncogenes and suppressor genes. Further, the predicted gene expressions could predict clinical phenotypes such as breast tumor status and estrogen receptor status (AUC=0.999, 0.94 respectively) as accurately as the measured gene expression levels. These results suggest that geneEXPLORER provides a means for accurate imputation of gene expression, which can be further used to predict clinical phenotypes.

bioinformatics

NaviCom: A web application to create interactive molecular network portraits using multi-level omics data

Human diseases such as cancer are routinely characterized by high-throughput molecular technologies, and multi-level omics data are accumulated in public databases at increasing rate. Retrieval and visualization of these data in the context of molecular network maps can provide insights into the pattern of molecular functions encompassed by an omics profile. In order to make this task easy, we developed NaviCom, a Python package and web platform for visualization of multi-level omics data on top of biological network maps. NaviCom is bridging the gap between cBioPortal, the most used resource of large-scale cancer omics data and NaviCell, a data visualization web service that contains several molecular network map collections. NaviCom proposes several standardized modes of data display on top of molecular network maps, allowing to address specific biological questions. We illustrate how users can easily create interactive network-based cancer molecular portraits via NaviCom web interface using the maps of Atlas of Cancer Signaling Network (ACSN) and other maps. Analysis of these molecular portraits can help in formulating a scientific hypothesis on the molecular mechanisms deregulated in the studied disease.

systems biology

Error-prone bypass of DNA lesions during lagging strand replication is a common source of germline and cancer mutations

Spontaneously occurring mutations are of great relevance in diverse fields including biochemistry, oncology, evolutionary biology, and human genetics. Studies in experimental systems have identified a multitude of mutational mechanisms including DNA replication infidelity as well as many forms of DNA damage followed by inefficient repair or replicative bypass. However, the relative contributions of these mechanisms to human germline mutations remain completely unknown. Here, based on the mutational asymmetry with respect to the direction of replication and transcription, we suggest that error-prone damage bypass on the lagging strand plays a major role in human mutagenesis. Asymmetry with respect to transcription is believed to be mediated by the action of transcription-coupled DNA repair (TC-NER). TC-NER selectively repairs DNA lesions on the transcribed strand; as a result, lesions on the non-transcribed strand are preferentially converted into mutations. In human polymorphism we detect a striking similarity between transcriptional asymmetry and asymmetry with respect to replication fork direction. This parallels the observation that damage-induced mutations in human cancers accumulate asymmetrically with respect to the direction of replication, suggesting that DNA lesions are asymmetrically resolved during replication. Re-analysis of XR-seq data, Damage-seq data and cancers with defective NER corroborate the preferential error-prone bypass of DNA lesions on the lagging strand. We experimentally demonstrate that replication delay greatly attenuates the mutagenic effect of UV-irradiation, in line with the key role of replication in conversion of DNA damage to mutations. We conservatively estimate that at least 10% of human germline mutations arise due to DNA damage rather than replication infidelity. The number of these damage-induced mutations is expected to scale with the number of replications and, consequently, paternal age.

biochemistry

Autogenous Cross-Regulation Of Quaking mRNA Processing And Translation Balances Quaking Functions In Splicing And Translation

Quaking RNA binding protein (RBP) isoforms arise from a single Quaking gene, and bind the same RNA motif to regulate splicing, stability, decay, and localization of a large set of RNAs. However, the mechanisms by which the expression of this single gene is controlled to distribute appropriate amounts of each Quaking isoform to regulate such disparate gene expression processes are unknown. Here we explore the separate mechanisms that regulate expression of two isoforms, Quaking-5 (Qk5) and Quaking-6 (Qk6), in mouse muscle cells. We first demonstrate that Qk5 and Qk6 proteins have distinct functions in splicing and translation respectively, enforced primarily through differential subcellular localization. Using isoform-specific depletion, we find both Qk5 and Qk6 act through cis and trans post-transcriptional regulatory mechanisms on their own and each others transcripts, creating a network of auto- and cross-regulatory controls. Qk5 has a major role in nuclear RNA stability and splicing, whereas Qk6 acts through translational regulation. In different cell types the cross-regulatory influences discovered here generate a spectrum of Qk5/Qk6 ratios subject to additional cell type and developmental controls. These unexpectedly complex feedback loops underscore the importance of the balance of Qk isoforms, especially where they are key regulators of development and cancer.

molecular biology

Competing paths over fitness valleys in growing populations

Investigating the emergence of a particular cell type is a recurring theme in models of growing cellular populations. The evolution of resistance to therapy is a classic example. Common questions are: when does the cell type first occur, and via which sequence of steps is it most likely to emerge? For growing populations, these questions can be formulated in a general framework of branching processes spreading through a graph from a root to a target vertex. Cells have a particular fitness value on each vertex and can transition along edges at specific rates. Vertices represents cell states, say genotypes or physical locations, while possible transitions are acquiring a mutation or cell migration. We focus on the setting where cells at the root vertex have the highest fitness and transition rates are small. Simple formulas are derived for the time to reach the target vertex and for the probability that it is reached along a given path in the graph. We demonstrate our results on several scenarios relevant to the emergence of drug resistance, including: the orderings of resistance-conferring mutations in bacteria and the impact of imperfect drug penetration in cancer.

evolutionary biology

A Biomaterial Screening Approach to Reveal Microenvironmental Mechanisms of Drug Resistance

TOC FigureDrug response screening, gene expression, and kinome signaling were combined across biomaterial platforms to combat adaptive resistance to sorafenib.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=158 SRC=\"FIGDIR/small/168039_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (60K):\norg.highwire.dtl.DTLVardef@a7af0eorg.highwire.dtl.DTLVardef@d5f2bborg.highwire.dtl.DTLVardef@330bd6org.highwire.dtl.DTLVardef@14e1f1d_HPS_FORMAT_FIGEXP M_FIG C_FIG Insight BoxWe combined biomaterial platforms, drug screening, and systems biology to identify mechanisms of extracellular matrix-mediated adaptive resistance to RTK-targeted cancer therapies. Drug response was significantly varied across biomaterials with altered stiffness, dimensionality, and cell-cell contacts, and kinome reprogramming was responsible for these differences in drug sensitivity. Screening across many platforms and applying a systems biology analysis were necessary to identify MEK phosphorylation as the key factor associated with variation in drug response. This method uncovered the combination therapy of sorafenib with a MEK inhibitor, which decreased viability on and within biomaterials in vitro, but was not captured by screening on tissue culture plastic alone. This combination therapy also reduced tumor burden in vivo, and revealed a promising approach for combating adaptive drug resistance.\n\nAbstractTraditional drug screening methods lack features of the tumor microenvironment that contribute to resistance. Most studies examine cell response in a single biomaterial platform in depth, leaving a gap in understanding how extracellular signals such as stiffness, dimensionality, and cell-cell contacts act independently or are integrated within a cell to affect either drug sensitivity or resistance. This is critically important, as adaptive resistance is mediated, at least in part, by the extracellular matrix (ECM) of the tumor microenvironment. We developed an approach to screen drug responses in cells cultured on 2D and in 3D biomaterial environments to explore how key features of ECM mediate drug response. This approach uncovered that cells on 2D hydrogels and spheroids encapsulated in 3D hydrogels were less responsive to receptor tyrosine kinase (RTK)-targeting drugs sorafenib and lapatinib, but not cytotoxic drugs, compared to single cells in hydrogels and cells on plastic. We found that transcriptomic differences between these in vitro models and tumor xenografts did not reveal mechanisms of ECM-mediated resistance to sorafenib. However, a systems biology analysis of phospho-kinome data uncovered that variation in MEK phosphorylation was associated with RTK-targeted drug resistance. Using sorafenib as a model drug, we found that co-administration with a MEK inhibitor decreased ECM-mediated resistance in vitro and reduced in vivo tumor burden compared to sorafenib alone. In sum, we provide a novel strategy for identifying and overcoming ECM-mediated resistance mechanisms by performing drug screening, phospho-kinome analysis, and systems biology across multiple biomaterial environments.

bioengineering

Fragmentation of pooled PCR products for highly multiplexed TILLING

Improvements to massively parallel sequencing have allowed the routine recovery of natural and induced sequence variants. A broad range of biological disciplines have benefited from this, ranging from plant breeding to cancer research. The need for high sequence coverage to accurately recover single nucleotide variants and small insertions and deletions limits the applicability of whole genome approaches. This is especially true in organisms with a large genome size or for applications requiring the screening of thousands of individuals, such as the reverse-genetic technique known as TILLING, where whole genome approaches become cost prohibitive. Using PCR to target and sequence chosen genomic regions provides an attractive alternative as the vast reduction in interrogated bases means that sample size can be dramatically increased through amplicon multiplexing and multidimensional sample pooling while maintaining suitable coverage for recovery of small mutations. Direct sequencing of PCR products is limited, however, due to limitations in read lengths of many next generation sequencers. In the present study we show the optimization and use of ultrasonication for the simultaneous fragmentation of multiplexed PCR amplicons for TILLING highly pooled samples. Thirty-two PCR products were produced from genomic DNA pools representing 265 pooled barley mutant lines. Mutant lines were produced with the chemical mutagen ethyl methanesulfonate. Samples were subjected to 2x300PE illumina sequencing. Evaluation of read coverage and base quality across amplicons suggests this approach is suitable for high-throughput TILLING and other applications employing highly pooled complex sampling schemes. Induced mutations previously identified in a traditional TILLING screen were recovered in this dataset further supporting the efficacy of the approach.

plant biology

EBADIMEX: An empirical Bayes approach to detect joint differential expression and methylation and to classify samples

1DNA methylation and gene expression are interdependent and both implicated in cancer development and progression, with many individual biomarkers discovered. A joint analysis of the two data types can potentially lead to biological insights that are not discoverable with separate analyses. To optimally leverage the joint data for identifying perturbed genes and classifying clinical cancer samples, it is important to accurately model the interactions between the two data types.\n\nHere, we present EBADIMEX for jointly identifying differential expression and methylation and classifying samples. The moderated t-test widely used with empirical Bayes priors in current differential expression methods is generalised to a multivariate setting by developing: (1) a moderated Welch t-test for equality of means with unequal variances; (2) a moderated F-test for equality of variances; and (3) a multivariate test for equality of means with equal variances. This leads to parametric models with prior distributions for the parameters, which allow fast evaluation and robust analysis of small data sets.\n\nEBADIMEX is demonstrated on simulated data as well as a large breast cancer (BRCA) cohort from TCGA. We show that the use of empirical Bayes priors and moderated tests works particularly well on small data sets.

bioinformatics

A network of epigenomic and transcriptional cooperation encompassing an epigenomic master regulator in cancer

Coordinated experiments focused on transcriptional responses and chromatin states are well-equipped to capture different epigenomic and transcriptomic levels governing the circuitry of a regulatory network. We propose a workflow for the genome-wide identification of epigenomic and transcriptional cooperation to elucidate transcriptional networks in cancer. Gene promoter annotation in combination with network analysis and sequence-resolution of enriched transcriptional motifs in epigenomic data reveals transcription factor families that act synergistically with epigenomic master regulators. A close teamwork of the transcriptional and epigenomic machinery was discovered. The network is tightly connected and includes the histone lysine demethylase KDM3A, basic helix-loop-helix factors MYC, HIF1A, and SREBF1, as well as differentiation factors AP1, MYOD1, SP1, MEIS1, ZEB1 and ELK1. In such a cooperation network, one component opens the chromatin, another one recognizes gene-specific DNA motifs, others scaffold between histones, cofactors, and the transcriptional complex. In cancer, due to the ability to team up with transcription factors, epigenetic factors concert mitogenic and metabolic gene networks, claiming the role of a cancer master regulators or epioncogenes.\n\nSpecific histone modification patterns are commonly associated with open or closed chromatin states, and are linked to distinct biological outcomes by transcriptional activation or repression. Disruption of patterns of histone modifications is associated with loss of proliferative control and cancer. There is tremendous therapeutic potential in understanding and targeting histone modification pathways. Thus, investigating cooperation of chromatin remodelers and the transcriptional machinery is not only important for elucidating fundamental mechanisms of chromatin regulation, but also necessary for the design of targeted therapeutics.

systems biology

Causal interactions from proteomic profiles: molecular data meets pathway knowledge

Measurement of changes in protein levels and in post-translational modifications, such as phosphorylation, can be highly informative about the phenotypic consequences of genetic differences or about the dynamics of cellular processes. Typically, such proteomic profiles are interpreted intuitively or by simple correlation analysis. Here, we present a computational method to generate causal explanations for proteomic profiles using prior mechanistic knowledge in the literature, as recorded in cellular pathway maps. To demonstrate its potential, we use this method to analyze the cascading events after EGF stimulation of a cell line, to discover new pathways in platelet activation, to identify influential regulators of oncoproteins in breast cancer, to describe signaling characteristics in predefined subtypes of ovarian and breast cancers, and to highlight which pathway relations are most frequently activated across 32 cancer types. Causal pathway analysis, that combines molecular profiles with prior biological knowledge captured in computational form, may become a powerful discovery tool as the amount and quality of cellular profiling rapidly expands. The method is freely available at http://causalpath.org.

systems biology

Combinatorial Detection of Conserved Alteration Patterns for Identifying Cancer Subnetworks

BackgroundAdvances in large scale tumor sequencing have lead to an understanding that there are combinations of genomic and transcriptomic alterations speciflc to tumor types, shared across many patients. Unfortunately, computational identiflcation of functionally meaningful shared alteration patterns, impacting gene/protein interaction subnetworks, has proven to be challenging. FindingsWe introduce a novel combinatorial method, cd-CAP, for simultaneous detection of connected subnetworks of an interaction network where genes exhibit conserved alteration patterns across tumor samples. Our method differentiates distinct alteration types associated with each gene (rather than relying on binary information of a gene being altered or not), and simultaneously detects multiple alteration proflle conserved subnetworks. ConclusionsIn a number of The Cancer Genome Atlas (TCGA) data sets, cd-CAP identifled large biologically signiflcant subnetworks with conserved alteration patterns, shared across many tumor samples.

systems biology

HeartBioPortal: an internet-of-omics for human cardiovascular disease data

Cardiovascular disease (CVD) is the leading cause of death worldwide, responsible for over 17 million deaths annually, a rate which outpaces even that related to cancer. Despite these sobering statistics, the state-of-the-art in computational infrastructure for the study of contemporary datasets related to CVD lags substantially behind that widely available in oncology, where improved data science and visualization methods have delivered publicly available comprehensive cancer genomics resources like Memorial Sloan Kettering Cancer Centers cBioPortal1,2 and the National Cancer Institutes Genomic Data Commons (GDC) Portal3,4. In our view, such portals do an outstanding job of transforming data from The Cancer Genome Atlas (TCGA) into logical data visualizations that provide additional biological insight. Developing a similar user-friendly computational platform for CVD could signi ...

bioinformatics

Combining Bayesian Approaches and Evolutionary Techniques for the Inference of Breast Cancer Networks

Gene and protein networks are very important to model complex large-scale systems in molecular biology. Inferring or reverseengineering such networks can be defined as the process of identifying gene/protein interactions from experimental data through computational analysis. However, this task is typically complicated by the enormously large scale of the unknowns in a rather small sample size. Furthermore, when the goal is to study causal relationships within the network, tools capable of overcoming the limitations of correlation networks are required. In this work, we make use of Bayesian Graphical Models to attach this problem and, specifically, we perform a comparative study of different state-of-the-art heuristics, analyzing their performance in inferring the structure of the Bayesian Network from breast cancer data.

bioinformatics

The trade-off between parsimony and model complexity for understanding biomedical mechanisms from mathematical models

Mechanistic mathematical models have been used extensively to provide a deeper understanding of biological mechanisms, including unveiling the regulation of tumour growth and its response to various treatments. However, given the breadth of biological regulatory mechanisms, these models are frequently large and thus prone to potential issues with parameter identifiability. Statistical metrics like the Akaike and Bayesian information criteria can help identify a parsimonious model by balancing goodness of fit against model complexity. Yet simple models may fail to provide sufficient biological insight if they do not adequately capture known physiological processes or mechanisms. A modeller must therefore balance hypothesis generation and biological learning with model tractability. Here, we illustrate this balance using models of ovarian cancer growth and treatment response to cisplatin and immune checkpoint blockade in homologous recombination (HR)-deficient and HR-proficient immunocompetent mouse models. We develop a hierarchy of mathematical models of increasing complexity to describe tumour growth, treatment response, and immune dynamics. Our results highlight the limits of relying purely on statistical metrics for model selection, particularly when the goal is to obtain biological insight and underscore the importance of balancing model complexity to avoid overfitting and parameter unidentifiability.

systems biology

Identifying Network Perturbation in Cancer

We present a computational framework, called DISCERN (DIfferential SparsE Regulatory Network), to identify informative topological changes in gene-regulator dependence networks inferred on the basis of mRNA expression datasets within distinct biological states. DISCERN takes two expression datasets as input: an expression dataset of diseased tissues from patients with a disease of interest and another expression dataset from matching normal tissues. DISCERN estimates the extent to which each gene is perturbed - having distinct regulator connectivity in the inferred gene-regulatory dependencies between the disease and normal conditions. This approach has distinct advantages over existing methods. First, DISCERN infers conditional dependencies between candidate regulators and genes, where conditional dependence relationships discriminate the evidence for direct interactions from indirect interactions more precisely than pairwise correlation. Second, DISCERN uses a new likelihood-based scoring function to alleviate concerns about accuracy of the specific edges inferred in a particular network. DISCERN identifies perturbed genes more accurately in synthetic data than existing methods to identify perturbed genes between distinct states. In expression datasets from patients with acute myeloid leukemia (AML), breast cancer and lung cancer, genes with high DISCERN scores in each cancer are enriched for known tumor drivers, genes associated with the biological processes known to be important in the disease, and genes associated with patient prognosis, in the respective cancer. Finally, we show that DISCERN can uncover potential mechanisms underlying network perturbation by explaining observed epigenomic activity patterns in cancer and normal tissue types more accurately than alternative methods, based on the available epigenomic from the ENCODE project.

Genomics

Global analysis of N6-methyladenosine functions and its disease association using deep learning and network-based methods

N6-methyladenosine (m6A) is the most abundant methylation, existing in >25% of human mRNAs. Exciting recent discoveries indicate the close involvement of m6A in regulating many different aspects of mRNA metabolism and diseases like cancer. However, our current knowledge about how m6A levels are controlled and whether and how regulation of m6A levels of a specific gene can play a role in cancer and other diseases is mostly elusive. We propose in this paper a computational scheme for predicting m6A-regulated genes and m6A-associated disease, which includes Deep-m6A, the first model for detecting condition-specific m6A sites from MeRIP-Seq data with a single base resolution using deep learning and a new network-based pipeline that prioritizes functional significant m6A genes and its associated diseases using the Protein-Protein Interaction (PPI) and gene-disease heterogeneous networks. We applied Deep-m6A and this pipeline to 75 MeRIP-seq human samples, which produced a compact set of 709 functionally significant m6A-regulated genes and nine functionally enriched subnetworks. The functional enrichment analysis of these genes and networks reveal that m6A targets key genes of many critical biological processes including transcription, cell organization and transport, and cell proliferation and cancer-related pathways such as Wnt pathway. The m6A-associated disease analysis prioritized five significantly associated diseases including leukemia and renal cell carcinoma. These results demonstrate the power of our proposed computational scheme and provide new leads for understanding m6A regulatory functions and its roles in diseases.\n\nAuthor summaryThe goal of this work is to identify functional significant m6A-regulated genes and m6A-associated diseases from analyzing an extensive collection of MeRIP-seq data. To achieve this, we first developed Deep-m6A, a CNN model for single-base m6A prediction. To our knowledge, this is the first condition-specific single-base m6A site prediction model that combines mRNA sequence feature and MeRIP-Seq data. The 10-fold cross-validation and test on an independent dataset showthat Deep-m6A outperformed two sequence-based models. We applied Deep-m6A followed by network-based analysis using HotNet2 and RWRH to 75 human MeRIP-Seq samples from various cells and tissue under different conditions to globally detect m6A-regulated genes and further predict m6A mediated functions and associated diseases. This is also to our knowledge the first attempt to predict m6A functions and associated diseases using only computational methods in a global manner on a large number of human MeRIP-Seq samples. The predicted functions and diseases show considerable consistent with those reported in the literature, which demonstrated the power of our proposed pipeline to predict potential m6A mediated functions and associated diseases.

bioinformatics

In Vivo Flow Cytometry of Extremely Rare Circulating Cells

Circulating tumor cells (CTCs) are of great interest in cancer research, but methods for their enumeration remain far from optimal. We developed a new small animal research tool called \"Diffuse in vivo Flow Cytometry\" (DiFC) for detecting extremely rare fluorescently-labeled circulating cells directly in the bloodstream. The technique exploits near-infrared diffuse photons to detect and count cells flowing in large superficial arteries and veins without drawing blood samples. DiFC uses custom-designed, dual fiber optic probes that are placed in contact with the skin surface approximately above a major vascular bundle. In combination with a novel signal processing, algorithm DiFC allows counting of individual cells moving in arterial or venous directions, as well as measurement of their speed and depth. We show that DiFC allows sampling of the entire circulating blood volume of a mouse in under 10 minutes, while maintaining a false alarm rate of 0.014 per minute. Hence, the unique capabilities of DiFC are highly suited to biological applications involving very rare cell types such as the study of hematogenic cancer metastasis.

bioengineering