Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Hierarchical organization of the human cell from a cancer coessentiality network

Genetic interactions mediate the emergence of phenotype from genotype. Systematic survey of genetic interactions in yeast showed that genes operating in the same biological process have highly correlated genetic interaction profiles, and this observation has been exploited to infer gene function in model organisms. Systematic surveys of digenic perturbations in human cells are also highly informative, but are not scalable, even with CRISPR-mediated methods. As an alternative, we developed an indirect method of deriving functional interactions. We show that genes having correlated knockout fitness profiles across diverse, non-isogenic cell lines are analogous to genes having correlated genetic interaction profiles across isogenic query strains, and similarly implies shared biological function. We constructed a network of genes with correlated fitness profiles across 400 CRISPR knockout screens in cancer cell lines into a \"coessentiality network,\" with up to 500-fold enrichment for co-functional gene pairs, enabling strong inference of human gene function. Modules in the network are connected in a layered web that gives insight into the hierarchical organization of the cell.

systems biology

Connecting tumor genomics with therapeutics through multi-dimensional network modules

Recent efforts have catalogued genomic, transcriptomic, epigenetic and proteomic changes in tumors, but connecting these data with effective therapeutics remains a challenge. In contrast, cancer cell lines can model therapeutic responses but only partially reflect tumor biology. Bridging this gap requires new methods of data integration to identify a common set of pathways and molecular events. Using MAGNETIC, a new method to integrate molecular profiling data using functional networks, we identify 219 gene modules in TCGA breast cancers that capture recurrent alterations, reveal new roles for H3K27 tri-methylation and accurately quantitate various cell types within the tumor microenvironment. We show that a significant portion of gene expression and methylation in tumors is poorly reproduced in cell lines due to differences in biology and microenvironment and MAGNETIC identifies therapeutic biomarkers that are robust to these differences. This work addresses a fundamental challenge in pharmacogenomics that can only be overcome by the joint analysis of patient and cell line data.

bioinformatics

MSIGNET: a Metropolis sampling-based method for global optimal significant network identification

In this paper, we propose a novel approach namely MSIGNET to identify subnetworks with significantly expressed genes by integrating context specific gene expression and protein-protein interaction (PPI) data. Specifically, we integrate differential expression of each gene and mutual information of gene pairs in a Bayesian framework and use Metropolis sampling to identify functional interactions. During the sampling process, a conditional probability is calculated given a randomly selected gene to control the network state transition. Our method provides global statistics of all genes and their interactions, and finally achieves a global optimal sub-network. We apply MSIGNET to simulated data and have demonstrated its superior performance over comparable network identification tools. Using a validated Parkinson data set we show that the network identified using MSIGNET is consistent to previously reported results but provides more biology meaningful interpretation of Parkinsons disease. Finally, to study networks related to ovarian cancer recurrence, we investigate two patient data sets. Identified networks from independent data sets show functional consistence. And those common genes and interactions are well supported by current biological knowledge.

bioinformatics

D3M: Detection of differential distributions of methylation levels

Motivation: DNA methylation is an important epigenetic modification related to a variety of diseases including cancers. We focus on the methylation data from Illuminas Infinium HumanMethylation450 BeadChip. One of the key issues of methylation analysis is to detect the differential methylation sites between case and control groups. Previous approaches describe data with simple summary statistics and kernel function, and then use statistical tests to determine the difference. However, a summary statistics-based approach cannot capture complicated underlying structure, and a kernel functions-based approach lacks interpretability of results.\n\nResults: We propose a novel method D3M, for detection of differential distribution of methylation, based on distribution-valued data. Our method can detect high-order moments, such as shapes of underlying distributions in methylation profiles, based on the Wasserstein metric. We test the significance of the difference between case and control groups and provide an interpretable summary of the results. The simulation results show that the proposed method achieves promising accuracy and shows favorable results compared with previous methods. Glioblastoma multiforme and lower grade glioma data from The Cancer Genome Atlas show that our method supports recent biological advances and suggests new insights.\n\nAvailability: R implemented code is freely available from\n\nhttps://github.com/ymatts/D3M/\n\nhttps://cran.r-project.org/package=D3M.\n\nContact: ymatsui@med.nagoya-u.ac.jp

Bioinformatics

Robust and Stable Gene Selection via Maximum-Minimum Correntropy Criterion

One of the central challenges in cancer research is identifying significant genes among thousands of others on a microarray. Since preventing outbreak and progression of cancer is the ultimate goal in bioinformatics and computational biology, detection of genes that are most involved is vital and crucial. In this article, we propose a Maximum-Minimum Correntropy Criterion (MMCC) approach for selection of biologically meaningful genes from microarray data sets which is stable, fast and robust against diverse noise and outliers and competitively accurate in comparison with other algorithms. Moreover, via an evolutionary optimization process, the optimal number of features for each data set is determined. Through broad experimental evaluation, MMCC is proved to be significantly better compared to other well-known gene selection algorithms for 25 commonly used microarray data sets. Surprisingly, high accuracy in classification by Support Vector Machine (SVM) is achieved by less than10 genes selected by MMCC in all of the cases.

Bioinformatics

ChimPipe: Accurate detection of fusion genes and transcription-induced chimeras from RNA-seq data

BackgroundChimeric transcripts are commonly defined as transcripts linking two or more different genes in the genome, and can be explained by various biological mechanisms such as genomic rearrangement, read-through or trans-splicing, but also by technical or biological artefacts. Several studies have shown their importance in cancer, cell pluripotency and motility. Many programs have recently been developed to identify chimeras from Illumina RNA-seq data (mostly fusion genes in cancer). However outputs of different programs on the same dataset can be widely inconsistent, and tend to include many false positives. Other issues relate to simulated datasets restricted to fusion genes, real datasets with limited numbers of validated cases, result inconsistencies between simulated and real datasets, and gene rather than junction level assessment.\n\nResultsHere we present ChimPipe, a modular and easy-to-use method to reliably identify chimeras from paired-end Illumina RNA-seq data. We have also produced realistic simulated datasets for three different read lengths, and enhanced two gold-standard cancer datasets by associating exact junction points to validated gene fusions. Benchmarking ChimPipe together with four other state-of-the-art tools on this data showed ChimPipe to be the top program at identifying exact junction coordinates for both kinds of datasets, and the one showing the best trade-off between sensitivity and precision. Applied to 106 ENCODE human RNA-seq datasets, ChimPipe identified 137 high confidence chimeras connecting the protein coding sequence of their parent genes. In subsequent experiments, three out of four predicted chimeras, two of which recurrently expressed in a large majority of the samples, could be validated. Cloning and sequencing of the three cases revealed several new chimeric transcript structures, 3 of which with the potential to encode a chimeric protein for which we hypothesized a new role.\n\nConclusionsChimPipe combines spanning and paired end RNA-seq reads to detect any kind of chimeras, including read-throughs, and shows an excellent trade-off between sensitivity and precision. The chimeras found by ChimPipe can be validated in-vitro with high accuracy.

Bioinformatics

MAGIC: A diffusion-based imputation method reveals gene-gene interactions in single-cell RNA-sequencing data

Single-cell RNA-sequencing is fast becoming a major technology that is revolutionizing biological discovery in fields such as development, immunology and cancer. The ability to simultaneously measure thousands of genes at single cell resolution allows, among other prospects, for the possibility of learning gene regulatory networks at large scales. However, scRNA-seq technologies suffer from many sources of significant technical noise, the most prominent of which is dropout due to inefficient mRNA capture. This results in data that has a high degree of sparsity, with typically only ~10% non-zero values. To address this, we developed MAGIC (Markov Affinity-based Graph Imputation of Cells), a method for imputing missing values, and restoring the structure of the data. After MAGIC, we find that two- and three-dimensional gene interactions are restored and that MAGIC is able to impute complex and non-linear shapes of interactions. MAGIC also retains cluster structure, enhances cluster-specific gene interactions and restores trajectories, as demonstrated in mouse retinal bipolar cells, hematopoiesis, and our newly generated epithelial-to-mesenchymal transition dataset.

bioinformatics

clonealign: statistical integration of independent single-cell RNA & DNA-seq from human cancers

Measuring gene expression of genomically defined tumour clones at single cell resolution would associate functional consequences to somatic alterations, as a prelude to elucidating pathways driving cell population growth, resistance and relapse. In the absence of scalable methods to simultaneously assay DNA and RNA from the same single cell, independent sampling of cell populations for parallel measurement of single cell DNA and single cell RNA must be computationally mapped for genome-transcriptome association. Here we present clonealign, a robust statistical framework to assign gene expression states to cancer clones using single-cell RNA-seq and DNA-seq independently sampled from an heterogeneous cancer cell population. We apply clonealign to triple-negative breast cancer patient derived xenografts and high-grade serous ovarian cancer cell lines and discover clone-specific dysregulated biological pathways not visible using either DNA-Seq or RNA-Seq alone.

bioinformatics

Selective pressures on human cancer genes along the evolution of mammals

Cancer is a disease of the genome caused by somatic mutation and subsequent clonal selection. Several genes associated to cancer in humans, hereafter cancer genes, also show evidence of (germline) positive selection among species. Taking advantage of a large collection of mammalian genomes, we systematically looked for statistically significant signatures of positive selection using dN/dS models in a list of 430 cancer genes. Among these, we identified 63 genes under putative positive selection in mammals, which are significantly enriched in processes like crosslinking DNA repair. We also found evidence of a higher incidence of positive selection in cancer genes bearing germline mutations, like BRCA2, where positively selected residues are physically linked with known pathogenic variants, suggesting a potential association between germline positive selection and risk of hereditary cancer. Overall, our results suggest that genes associated with hereditary cancer have less selective constraints than genes related to sporadic cancer. Also, that the adaptive evolution of human cancer genes in mammals has been most likely driven by adaptive changes in important traits not directly related to cancer.

evolutionary biology

An Algorithm for Cellular Reprogramming

The day we understand the time evolution of subcellular elements at a level of detail comparable to physical systems governed by Newtons laws of motion seems far away. Even so, quantitative approaches to cellular dynamics add to our understanding of cell biology, providing data-guided frameworks that allow us to develop better predictions about, and methods for, control over specific biological processes and system-wide cell behavior. In this paper, we describe an approach to optimizing the use of transcription factors (TFs) in the context of cellular reprogramming. We construct an approximate model for the natural evolution of a cell cycle synchronized population of human fibroblasts, based on data obtained by sampling the expression of 22,083 genes at several time points along the cell cycle. In order to arrive at a model of moderate complexity, we cluster gene expression based on the division of the genome into topologically associating domains (TADs) and then model the dynamics of the TAD expression levels. Based on this dynamical model and known bioinformatics, such as transcription factor binding sites (TFBS) and functions, we develop a methodology for identifying the top transcription factor candidates for a specific cellular reprogramming task. The approach used is based on a device commonly used in optimal control. Our data-guided methodology identifies a number of transcription factors previously validated for reprogramming and/or natural differentiation. Our findings highlight the immense potential of dynamical models, mathematics, and data-guided methodologies for improving strategies for control over biological processes.\n\nSignificance StatementReprogramming the human genome toward any desirable state is within reach; application of select transcription factors drives cell types toward different lineages in many settings. We introduce the concept of data-guided control in building a universal algorithm for directly reprogramming any human cell type into any other type. Our algorithm is based on time series genome transcription and architecture data and known regulatory activities of transcription factors, with natural dimension reduction using genome architectural features. Our algorithm predicts known reprogramming factors, top candidates for new settings, and ideal timing for application of transcription factors. This framework can be used to develop strategies for tissue regeneration, cancer cell reprogramming, and control of dynamical systems beyond cell biology.

bioinformatics

netDx: Patient classification using integrated patient similarity networks

Patient classification has widespread biomedical and clinical applications, including diagnosis, prognosis and treatment response prediction. A clinically useful prediction algorithm should be accurate, generalizable, be able to integrate diverse data types, and handle sparse data. A clinical predictor based on genomic data needs to be easily interpretable to drive hypothesis-driven research into new treatments. We describe netDx, a novel supervised patient classification framework based on patient similarity networks. netDx meets the above criteria and particularly excels at data integration and model interpretability. As a machine learning method, netDx demonstrates consistently excellent performance in a cancer survival benchmark across four cancer types by integrating up to six genomic and clinical data types. In these tests, netDx has significantly higher average performance than most other machine-learning approaches across most cancer types and its best model outperforms all other methods for two cancer types. In comparison to traditional machine learning-based patient classifiers, netDx results are more interpretable, visualizing the decision boundary in the context of patient similarity space. When patient similarity is defined by pathway-level gene expression, netDx identifies biological pathways important for outcome prediction, as demonstrated in diverse data sets of breast cancer and asthma. Thus, netDx can serve both as a patient classifier and as a tool for discovery of biological features characteristic of disease. We provide a software complete implementation of netDx along with sample files and automation workflows in R.

bioinformatics

Semblance Of Heterogeneity In Collective Cell Migration

Cell population heterogeneity is increasingly a focus of inquiry in biological research. For example, cell migration studies have investigated the heterogeneity of invasiveness and taxis in development, wound healing, and cancer. However, relatively little effort has been devoted to explore when heterogeneity is mechanistically relevant and how to reliably measure it. Statistical methods from the animal movement literature offer the potential to analyse heterogeneity in collections of cell tracking data. A popular measure of heterogeneity, which we use here as an example, is the distribution of delays in directional cross-correlation. Using a suitably generic, yet minimal, model of collective cell movement in three dimensions, we show how using such measures to quantify heterogeneity in tracking data can result in the inference of heterogeneity where there is none. Our study highlights a potential pitfall in the statistical analysis of cell population heterogeneity, and we argue this can be mitigated by the appropriate choice of null models.\n\nHighlightsO_LIgroups of identical cells appear heterogeneous due to limited sampling and experimental repeatability\nC_LIO_LIheterogeneity bias increases with attraction/repulsion between cells\nC_LIO_LImovement in confined environments decreases apparent heterogeneity\nC_LIO_LIhypothetical applications in neural crest and in vitro cancer systems\nC_LI\n\nIn BriefWe use a mathematical model to show how cell populations can appear heterogeneous in their migratory characteristics, even though they are made up of identically-behaving individual cells. This has important consequences for the study of collective cell migration in areas such as embryo development or cancer invasion.

systems biology

Pathway-Structured Predictive Model for Cancer Survival Prediction: A Two-Stage Approach

Heterogeneity in terms of tumor characteristics, prognosis, and survival among cancer patients has been a persistent problem for many decades. Currently, prognosis and outcome predictions are made based on clinical factors and/or by incorporating molecular profiling data. However, inaccurate prognosis and prediction may result by using only clinical or molecular information directly. One of the main shortcomings of past studies is the failure to incorporate prior biological information into the predictive model, given strong evidence of pathway-based genetic nature of cancer, i.e. the potential for oncogenes to be grouped into pathways based on biological functions such as cell survival, proliferation and metastatic dissemination.\n\nTo address this problem, we propose a two-stage procedure to incorporate pathway information into the prognostic modeling using large-scale gene expression data. In the first stage, we fit all predictors within each pathway using penalized Cox model (Lasso, Ridge and Elastic Net) and Bayesian hierarchical Cox model. In the second stage, we combine the cross-validated prognostic scores of all pathways obtained in the first stage as new predictors to build an integrated prognostic model for prediction. We apply the proposed method to analyze breast cancer data from The Cancer Genome Atlas (TCGA), predicting overall survival using clinical data and gene expression profiling. The data includes ~20000 genes mapped into 109 pathways for 505 patients. The results show that the proposed approach not only improves survival prediction compared with the alternative analysis that ignores the pathway information, but also identifies significant biological pathways.

Bioinformatics

iCAGES: integrated CAncer GEnome Score for comprehensively prioritizing cancer driver genes in personal genomes

All cancers arise as a result of the acquisition of somatic mutations that drive the disease progression. A number of computational tools have been developed to identify driver genes for a specific cancer from a group of cancer samples. However, it remains a challenge to identify driver mutations/genes for an individual patient and design drug therapies. We developed iCAGES, a novel statistical framework to rapidly analyze patient-specific cancer genomic data, prioritize personalized cancer driver events and predict personalized therapies. iCAGES includes three consecutive layers: the first layer integrates contributions from coding, non-coding and structural variations to infer driver variants. For coding mutations, we developed a radial support vector machine using manually curated mutations to predict their driver potential. The second layer identifies driver genes, by using information from the first layer and integrating prior biological knowledge on gene-gene and gene-phenotype networks. The third layer prioritizes personalized drug treatment, by classifying potential driver genes into different categories and querying drug-gene databases. Compared to currently available tools, iCAGES achieves better performance by correctly classifying point coding driver mutations (AUC=0.97, 95% CI: 0.97-0.97, significantly better than the second best tool with P=0.01) and genes (AUC=0.93, 95% CI: 0.93-0.94, significantly better than MutSigCV with P<1x10-15). We also illustrated two examples where iCAGES correctly nominated two targeted drugs for two advanced cancer patients with exceptional response, based on their somatic mutation profiles. iCAGES leverages personal genomic information and prior biological knowledge, effectively identifies cancer driver genes and predicts treatment strategies. iCAGES is available at http://icages.usc.edu.

Genomics

A Standard Operating Procedure For Outlier Removal In Large-Sample Epidemiological Transcriptomics Datasets

Transcriptome measurements and other -omics type data are increasingly more used in epidemiological studies. Most of omics studies to date are small with samples sizes in the tens, or sometimes low hundreds, but this is changing. Our Norwegian Woman and Cancer (NOWAC) datasets are to date one or two orders of magnitude larger. The NOWAC biobank contains about 50000 blood samples from a prospective study. Around 125 breast cancer cases occur in this cohort each year. The large biological variation in gene expression means that many observations are needed to draw scientific conclusions. This is true for both microarray and RNA-seq type data. Hence, larger datasets are likely to become more common soon.\n\nTechnical outliers are observations that somehow were distorted at the lab or during sampling. If not removed these observations add bias and variance in later statistical analyses, and may skew the results. Hence, quality assessment and data cleaning are important. We find common quality assessment libraries difficult to work with for large datasets for two reasons: slow execution speed and unsuitable visualizations.\n\nIn this paper, we present our standard operating procedure (SOP) for large-sample transcriptomics datasets. Our SOP combines automatic outlier detection with manual evaluation to avoid removing valuable observations. We use laboratory quality measures and statistical measures of deviation to aid the analyst. These are available in the nowaclean R package, currently available on GitHub (https://github.com/3inar/nowaclean). Finally, we evaluate our SOP on one of our larger datasets with 832 observations.

epidemiology

Anti-mutagenic and synergistic cytotoxic effect of cisplatin and Honey Bee venom on 4T1 invasive mammary carcinoma cell line

Honey Bee Venom has various biological activities such as inhibitory effect on several types of cancer. Cisplatin is an old and potent drug to treat the most of cancer. Our aims in this study were determination of the anti-mutagenic and cytotoxic effects of HBV on mammary carcinoma, lonely and in combination with cisplatin. In this study 4T1 cell line were cultured and incubated at 37 C in humidified CO2-incubator. The cell viabilities were examined by MTT assay. Also HBV was screened for its anti-mutagenic activity against sodium azide by Ames test. The result showed that 6g/ml HBV, 20g/ml cisplatin and 6g/ml HBV with 10g/ml cisplatin can induce an approximately 50% 4T1 cell death. 7mg/ml HBV with the inhibition of 62.76% sodium azide showed high potential in decreasing the mutagenic agents. MTT assay demonstrated that HBV and cisplatin can cause cell death in a dose-dependent manner. The cytotoxic effect of cisplatin is also promoted by HBV. Ames test results indicated that HBV can inhibit sodium azide as a mutagenic agent. Anti-mutagenic activity of HBV was increased significantly in presence of S9 mix. Hence, our findings reveal that HBV can enhance the cytotoxic effect of cisplatin drug and it has cancer preventing effects.

pharmacology and toxicology

Regulation of cancer epigenomes with a histone-binding synthetic transcription factor

Chromatin proteins have expanded the mammalian synthetic biology toolbox by enabling control of active and silenced states at endogenous genes. Others have reported synthetic proteins that bind DNA and regulate genes by altering chromatin marks, such as histone modifications. Previously we reported the first synthetic transcriptional activator, the \"Polycomb-based transcription factor\" (PcTF), that reads histone modifications through a protein-protein interaction between the PCD motif and trimethylated lysine 27 of histone H3 (H3K27me3). Here, we describe the genome-wide behavior of PcTF. Transcriptome and chromatin profiling revealed PcTF-sensitive promoter regions marked by proximal PcTF and distal H3K27me3 binding. These results illuminate a mechanism in which PcTF interactions bridge epigenetic marks with the transcription initiation complex. In three cancer-derived human cell lines tested here, many PcTF-sensitive genes encode developmental regulators and tumor suppressors. Thus, PcTF represents a powerful new fusion-protein-based method for cancer research and treatment where silencing marks are translated into direct gene activation.

Synthetic Biology

Allosteric activation dictates PRC2 activity independent of its recruitment to chromatin

PRC2 is a therapeutic target for several types of cancers currently undergoing clinical trials. Its activity is regulated by a positive feedback loop whereby its terminal enzymatic product, H3K27me3, is specifically recognized and bound by an aromatic cage present in its EED subunit. The ensuing allosteric activation of the complex stimulates H3K27me3 deposition on chromatin. Here, we report a step-wise feedback mechanism entailing key residues within distinctive interfacing motifs of EZH2 or EED that are found mutated in cancers and/or Weaver syndrome. PRC2 harboring these EZH2 or EED mutants manifest little activity in vivo but, unexpectedly, exhibited similar chromatin association as wild-type PRC2, indicating an uncoupling of PRC2 activity and recruitment. With genetic and chemical tools, we further demonstrated that targeting allosteric activation overrode the gain-of-function effect of EZH2Y646X oncogenic mutations. These results revealed critical implications to the regulation and biology of PRC2 and a novel vulnerability in tackling PRC2-addicted cancers.

biochemistry