Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,801 records · Page 100Linked to original sources

Detection and quantification of viral RNA in human tumors using open source pipeline: viGEN

An estimated 17% of cancers worldwide are associated with infectious causes. The extent and biological significance of viral presence/infection in actual tumor samples is generally unknown but could be measured using human transcriptome (RNA-seq) data from tumor samples.\n\nWe present an open source bioinformatics pipeline viGEN, which combines existing well-known and novel RNA-seq tools for not only the detection and quantification of viral RNA, but also variants in the viral transcripts.\n\nThe pipeline includes 4 major modules: The first module allows to align and filter out human RNA sequences; the second module maps and count (remaining un-aligned) reads against reference genomes of all known and sequenced human viruses; the third module quantifies read counts at the individual viral genes level thus allowing for downstream differential expression analysis of viral genes between experimental and controls groups. The fourth module calls variants in these viruses. To the best of our knowledge, there are no publicly available pipelines or packages that would provide this type of complete analysis in one open source package.\n\nIn this paper, we applied the viGEN pipeline to two case studies. We first demonstrate the working of our pipeline on a large public dataset, the TCGA cervical cancer cohort. We also performed additional in-depth analyses on a small focused study of TCGA liver cancer patients. In this cohort, we perform viral-gene quantification, viral-variant extraction and survival analysis. This allowed us to find differentially expressed viral-transcripts and viral-variants between the groups of patients, and connect them to clinical outcome.\n\nFrom our analyses, we show that we were able to successfully detect the human papilloma virus among the TCGA cervical cancer patients. We compared the viGEN pipeline with two metagenomics tools and demonstrate similar sensitivity/specificity. We were also able to quantify viral-transcripts and extract viral-variants using the liver cancer dataset. The results presented corresponded with published literature in terms of rate of detection, viral gene expression patterns and impact of several known variants of HBV genome. Results also show novel information about distinct patterns of expression and co-expression in Hepatitis B and the Human Endogenous Retrovirus (HERV) K113 viruses.\n\nThis pipeline is generalizable, and can be used to provide novel biological insights into the significance of viral and other microbial infections in complex diseases, tumorigeneses and cancer immunology. The source code, with example data and tutorial is available at: https://github.com/ICBI/viGEN/.

bioinformatics

Engineered probiotics for local tumor delivery of checkpoint blockade nanobodies

Immunotherapies such as checkpoint inhibitors have revolutionized cancer therapy yet lead to a multitude of immune-related adverse events, suggesting the need for more targeted delivery systems. Due to their preferential colonization of tumors and advances in engineering capabilities from synthetic biology, microbes are a natural platform for the local delivery of cancer therapeutics. Here, we present an engineered probiotic bacteria system for the controlled production and release of novel immune checkpoint targeting nanobodies from within tumors. Specifically, we engineered genetic lysis circuit variants to effectively release nanobodies and safely control bacteria populations. To maximize therapeutic efficacy of the system, we used computational modeling coupled with experimental validation of circuit dynamics and found that lower copy number variants provide optimal nanobody release. Thus, we subsequently integrated the lysis circuit operon into the genome of a probiotic E. coli Nissle 1917, and confirmed lysis dynamics in a syngeneic mouse model using in vivo bioluminescent imaging. Expressing a nanobody against PD-L1 in this strain demonstrated enhanced efficacy compared to a plasmid-based lysing variant, and similar efficacy to a clinically relevant monoclonal antibody against PD-L1. Expanding upon this therapeutic platform, we produced a nanobody against cytotoxic T-lymphocyte associated protein -4 (CTLA-4), which reduced growth rate or completely cleared tumors when combined with a probiotically-expressed PD-L1 nanobody in multiple syngeneic mouse models. Together, these results demonstrate that our engineered probiotic system combines innovations in synthetic biology and immunotherapy to improve upon the delivery of checkpoint inhibitors. SENTENCE SUMMARYWe designed a probiotic platform to locally deliver checkpoint blockade nanobodies to tumors using a controlled lysing mechanism for therapeutic release.

synthetic biology

Identifying miRNA-mRNA regulatory relationships in breast cancer with invariant causal prediction

microRNAs (miRNAs) regulate gene expression at the post-transcriptional level and they play an important role in various biological processes in the human body. Therefore, identifying their regulation mechanisms is essential for the diagnostics and therapeutics for a wide range of diseases. There have been a large number of researches which use gene expression profiles to resolve this problem. However, the current methods have their own limitations. Some of them only identify the correlation of miRNA and mRNA expression levels instead of the causal or regulatory relationships while others infer the causality but with a high computational complexity. To overcome these issues, in this study, we propose a method to identify miRNA-mRNA regulatory relationships in breast cancer using the invariant causal prediction. The key idea of invariant causal prediction is that the cause miRNAs of their target mRNAs are the ones which have persistent causal relationships with the target mRNAs across different environments. In this research, we aim to find miRNA targets which are consistent across different breast cancer subtypes. Thus, first of all, we apply the Pam50 method to categorise BRCA samples into different environment\" groups based on different cancer subtypes. Then we use the invariant causal prediction method to find miRNA-mRNA regulatory relationships across subtypes. We validate the results with the miRNA-transfected experimental data and the results show that our method outperforms the state-of-the-art methods. In addition, we also integrate this new method with the Pearson correlation analysis method and Lasso in an ensemble method to take the advantages of these methods. We then validate the results of the ensemble method with the experimentally confirmed data and the ensemble method shows the best performance, even comparing to the proposed causal method. Functional enrichment analyses show that miRNAs in the regulatory relationship predicated by the proposed causal method tend to synergistically regulate target genes, indicating the usefulness of these methods, and the identified miRNA targets could be used in the design of wet-lab experiments to discover the causes of breast cancer.\n\nAuthor summaryCancer is a disease of cells in human body and it causes a high rate of deaths world wide. There has been evidence that non-coding RNAs are key players in the development and progression of cancer. Among the different types of non-coding RNAs, miRNAs, which are short non-coding RNAs, regulate gene expression and play an important role in different biological processes as well as various cancer types. To design better diagnostic and therapeutic plans for cancer patients, we need to know the roles of miRNAs in cancer initialisation and development, and their regulation mechanisms in the human body. In this study, we propose algorithms to identify miRNA-mRNA regulatory relationships in breast cancer. Comparing our methods with existing methods in predicting miRNA targets, our methods show a better performance. The estimated miRNA targets from our methods could be a potential source for further wet-lab experiments to discover the causes of breast cancer.

bioinformatics

Memory sequencing reveals heritable single cell gene expression programs associated with distinct cellular behaviors

Non-genetic factors can cause individual cells to fluctuate substantially in gene expression levels over time. Yet it remains unclear whether these fluctuations can persist for much longer than the time of one cell division. Current methods for measuring gene expression in single cells mostly rely on single time point measurements, making the duration of gene expression fluctuations or cellular memory difficult to measure. Here, we report a method combining Luria and Delbrucks fluctuation analysis with population-based RNA sequencing (MemorySeq) for identifying genes transcriptome-wide whose fluctuations persist for several cell divisions. MemorySeq revealed multiple gene modules that are expressed together in rare cells within otherwise homogeneous clonal populations. Further, we found that these rare cell subpopulations are associated with biologically distinct behaviors, such as the ability to proliferate in the face of anti-cancer therapeutics, in different cancer cell lines. The identification of non-genetic, multigenerational fluctuations has the potential to reveal new forms of biological memory at the level of single cells and suggests that non-genetic heritability of cellular state may be a quantitative property.

systems biology

Addressing current challenges in cancer immunotherapy with mathematical and computational modeling

The goal of cancer immunotherapy is to boost a patients immune response to a tumor. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumor microenvironment, immune-modulating effects of conventional treatments, and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modeling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumor classification, optimal treatment scheduling, and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modelers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumor-immune biology. We conclude the review with recommendations for modelers both with respect to methodology and biological direction that might help keep modelers at the forefront of cancer immunotherapy development.

systems biology

TASI: A software tool for spatial-temporal quantification of tumor spheroid dynamics

Spheroid cultures derived from explanted cancer specimens are an increasingly utilized resource for studying complex biological processes like tumor cell invasion and metastasis, representing an important bridge between the simplicity and practicality of 2D monolayer cultures and the complexity and realism of in vivo animal models. Temporal imaging of spheroids can capture the dynamics of cell behaviors and microenvironments, and when combined with quantitative image analysis methods, enables deep interrogation of biological mechanisms. This paper presents a comprehensive open-source software framework for Temporal Analysis of Spheroid Imaging (TASI) that allows investigators to objectively characterize spheroid growth and invasion dynamics. TASI performs spatiotemporal segmentation of spheroid cultures, extraction of features describing spheroid morpho-phenotypes, mathematical modeling of spheroid dynamics, and statistical comparisons of experimental conditions. We demonstrate the utility of this tool in an analysis of non-small cell lung cancer spheroids that exhibit variability in metastatic and proliferative behaviors.

bioinformatics

Hierarchical organization of the human cell from a cancer coessentiality network

Genetic interactions mediate the emergence of phenotype from genotype. Systematic survey of genetic interactions in yeast showed that genes operating in the same biological process have highly correlated genetic interaction profiles, and this observation has been exploited to infer gene function in model organisms. Systematic surveys of digenic perturbations in human cells are also highly informative, but are not scalable, even with CRISPR-mediated methods. As an alternative, we developed an indirect method of deriving functional interactions. We show that genes having correlated knockout fitness profiles across diverse, non-isogenic cell lines are analogous to genes having correlated genetic interaction profiles across isogenic query strains, and similarly implies shared biological function. We constructed a network of genes with correlated fitness profiles across 400 CRISPR knockout screens in cancer cell lines into a \"coessentiality network,\" with up to 500-fold enrichment for co-functional gene pairs, enabling strong inference of human gene function. Modules in the network are connected in a layered web that gives insight into the hierarchical organization of the cell.

systems biology

Connecting tumor genomics with therapeutics through multi-dimensional network modules

Recent efforts have catalogued genomic, transcriptomic, epigenetic and proteomic changes in tumors, but connecting these data with effective therapeutics remains a challenge. In contrast, cancer cell lines can model therapeutic responses but only partially reflect tumor biology. Bridging this gap requires new methods of data integration to identify a common set of pathways and molecular events. Using MAGNETIC, a new method to integrate molecular profiling data using functional networks, we identify 219 gene modules in TCGA breast cancers that capture recurrent alterations, reveal new roles for H3K27 tri-methylation and accurately quantitate various cell types within the tumor microenvironment. We show that a significant portion of gene expression and methylation in tumors is poorly reproduced in cell lines due to differences in biology and microenvironment and MAGNETIC identifies therapeutic biomarkers that are robust to these differences. This work addresses a fundamental challenge in pharmacogenomics that can only be overcome by the joint analysis of patient and cell line data.

bioinformatics

MSIGNET: a Metropolis sampling-based method for global optimal significant network identification

In this paper, we propose a novel approach namely MSIGNET to identify subnetworks with significantly expressed genes by integrating context specific gene expression and protein-protein interaction (PPI) data. Specifically, we integrate differential expression of each gene and mutual information of gene pairs in a Bayesian framework and use Metropolis sampling to identify functional interactions. During the sampling process, a conditional probability is calculated given a randomly selected gene to control the network state transition. Our method provides global statistics of all genes and their interactions, and finally achieves a global optimal sub-network. We apply MSIGNET to simulated data and have demonstrated its superior performance over comparable network identification tools. Using a validated Parkinson data set we show that the network identified using MSIGNET is consistent to previously reported results but provides more biology meaningful interpretation of Parkinsons disease. Finally, to study networks related to ovarian cancer recurrence, we investigate two patient data sets. Identified networks from independent data sets show functional consistence. And those common genes and interactions are well supported by current biological knowledge.

bioinformatics

Chromatin interactome mapping at 139 independent breast cancer risk signals

Genome-wide association studies have identified 196 high confidence independent signals associated with breast cancer susceptibility. Variants within these signals frequently fall in distal regulatory DNA elements that control gene expression. We designed a Capture Hi-C array to enrich for chromatin interactions between the credible causal variants and target genes in six human mammary epithelial and breast cancer cell lines. We show that interacting regions are enriched for open chromatin, histone marks for active enhancers and transcription factors relevant to breast biology. We exploit this comprehensive resource to identify candidate target genes at 139 independent breast cancer risk signals, and explore the functional mechanism underlying altered risk at the 12q24 risk region. Our results demonstrate the power of combining genetics, computational genomics and molecular studies to rationalize the identification of key variants and candidate target genes at breast cancer GWAS signals.

molecular biology

D3M: Detection of differential distributions of methylation levels

Motivation: DNA methylation is an important epigenetic modification related to a variety of diseases including cancers. We focus on the methylation data from Illuminas Infinium HumanMethylation450 BeadChip. One of the key issues of methylation analysis is to detect the differential methylation sites between case and control groups. Previous approaches describe data with simple summary statistics and kernel function, and then use statistical tests to determine the difference. However, a summary statistics-based approach cannot capture complicated underlying structure, and a kernel functions-based approach lacks interpretability of results.\n\nResults: We propose a novel method D3M, for detection of differential distribution of methylation, based on distribution-valued data. Our method can detect high-order moments, such as shapes of underlying distributions in methylation profiles, based on the Wasserstein metric. We test the significance of the difference between case and control groups and provide an interpretable summary of the results. The simulation results show that the proposed method achieves promising accuracy and shows favorable results compared with previous methods. Glioblastoma multiforme and lower grade glioma data from The Cancer Genome Atlas show that our method supports recent biological advances and suggests new insights.\n\nAvailability: R implemented code is freely available from\n\nhttps://github.com/ymatts/D3M/\n\nhttps://cran.r-project.org/package=D3M.\n\nContact: ymatsui@med.nagoya-u.ac.jp

Bioinformatics

Robust and Stable Gene Selection via Maximum-Minimum Correntropy Criterion

One of the central challenges in cancer research is identifying significant genes among thousands of others on a microarray. Since preventing outbreak and progression of cancer is the ultimate goal in bioinformatics and computational biology, detection of genes that are most involved is vital and crucial. In this article, we propose a Maximum-Minimum Correntropy Criterion (MMCC) approach for selection of biologically meaningful genes from microarray data sets which is stable, fast and robust against diverse noise and outliers and competitively accurate in comparison with other algorithms. Moreover, via an evolutionary optimization process, the optimal number of features for each data set is determined. Through broad experimental evaluation, MMCC is proved to be significantly better compared to other well-known gene selection algorithms for 25 commonly used microarray data sets. Surprisingly, high accuracy in classification by Support Vector Machine (SVM) is achieved by less than10 genes selected by MMCC in all of the cases.

Bioinformatics

ChimPipe: Accurate detection of fusion genes and transcription-induced chimeras from RNA-seq data

BackgroundChimeric transcripts are commonly defined as transcripts linking two or more different genes in the genome, and can be explained by various biological mechanisms such as genomic rearrangement, read-through or trans-splicing, but also by technical or biological artefacts. Several studies have shown their importance in cancer, cell pluripotency and motility. Many programs have recently been developed to identify chimeras from Illumina RNA-seq data (mostly fusion genes in cancer). However outputs of different programs on the same dataset can be widely inconsistent, and tend to include many false positives. Other issues relate to simulated datasets restricted to fusion genes, real datasets with limited numbers of validated cases, result inconsistencies between simulated and real datasets, and gene rather than junction level assessment.\n\nResultsHere we present ChimPipe, a modular and easy-to-use method to reliably identify chimeras from paired-end Illumina RNA-seq data. We have also produced realistic simulated datasets for three different read lengths, and enhanced two gold-standard cancer datasets by associating exact junction points to validated gene fusions. Benchmarking ChimPipe together with four other state-of-the-art tools on this data showed ChimPipe to be the top program at identifying exact junction coordinates for both kinds of datasets, and the one showing the best trade-off between sensitivity and precision. Applied to 106 ENCODE human RNA-seq datasets, ChimPipe identified 137 high confidence chimeras connecting the protein coding sequence of their parent genes. In subsequent experiments, three out of four predicted chimeras, two of which recurrently expressed in a large majority of the samples, could be validated. Cloning and sequencing of the three cases revealed several new chimeric transcript structures, 3 of which with the potential to encode a chimeric protein for which we hypothesized a new role.\n\nConclusionsChimPipe combines spanning and paired end RNA-seq reads to detect any kind of chimeras, including read-throughs, and shows an excellent trade-off between sensitivity and precision. The chimeras found by ChimPipe can be validated in-vitro with high accuracy.

Bioinformatics

MAGIC: A diffusion-based imputation method reveals gene-gene interactions in single-cell RNA-sequencing data

Single-cell RNA-sequencing is fast becoming a major technology that is revolutionizing biological discovery in fields such as development, immunology and cancer. The ability to simultaneously measure thousands of genes at single cell resolution allows, among other prospects, for the possibility of learning gene regulatory networks at large scales. However, scRNA-seq technologies suffer from many sources of significant technical noise, the most prominent of which is dropout due to inefficient mRNA capture. This results in data that has a high degree of sparsity, with typically only ~10% non-zero values. To address this, we developed MAGIC (Markov Affinity-based Graph Imputation of Cells), a method for imputing missing values, and restoring the structure of the data. After MAGIC, we find that two- and three-dimensional gene interactions are restored and that MAGIC is able to impute complex and non-linear shapes of interactions. MAGIC also retains cluster structure, enhances cluster-specific gene interactions and restores trajectories, as demonstrated in mouse retinal bipolar cells, hematopoiesis, and our newly generated epithelial-to-mesenchymal transition dataset.

bioinformatics

clonealign: statistical integration of independent single-cell RNA & DNA-seq from human cancers

Measuring gene expression of genomically defined tumour clones at single cell resolution would associate functional consequences to somatic alterations, as a prelude to elucidating pathways driving cell population growth, resistance and relapse. In the absence of scalable methods to simultaneously assay DNA and RNA from the same single cell, independent sampling of cell populations for parallel measurement of single cell DNA and single cell RNA must be computationally mapped for genome-transcriptome association. Here we present clonealign, a robust statistical framework to assign gene expression states to cancer clones using single-cell RNA-seq and DNA-seq independently sampled from an heterogeneous cancer cell population. We apply clonealign to triple-negative breast cancer patient derived xenografts and high-grade serous ovarian cancer cell lines and discover clone-specific dysregulated biological pathways not visible using either DNA-Seq or RNA-Seq alone.

bioinformatics

Selective pressures on human cancer genes along the evolution of mammals

Cancer is a disease of the genome caused by somatic mutation and subsequent clonal selection. Several genes associated to cancer in humans, hereafter cancer genes, also show evidence of (germline) positive selection among species. Taking advantage of a large collection of mammalian genomes, we systematically looked for statistically significant signatures of positive selection using dN/dS models in a list of 430 cancer genes. Among these, we identified 63 genes under putative positive selection in mammals, which are significantly enriched in processes like crosslinking DNA repair. We also found evidence of a higher incidence of positive selection in cancer genes bearing germline mutations, like BRCA2, where positively selected residues are physically linked with known pathogenic variants, suggesting a potential association between germline positive selection and risk of hereditary cancer. Overall, our results suggest that genes associated with hereditary cancer have less selective constraints than genes related to sporadic cancer. Also, that the adaptive evolution of human cancer genes in mammals has been most likely driven by adaptive changes in important traits not directly related to cancer.

evolutionary biology

Unearthing Regulatory Axes of Breast Cancer circRNAs Networks to Find Novel Targets and Fathom Pivotal Mechanisms

Circular RNAs (circRNAs) along other complementary regulatory elements in ceRNAs networks possess valuable characteristics for both diagnosis and treatment of several human cancers including breast cancer (BC). In this study, we combined several systems biology tools and approaches to identify influential BC circRNAs, RNA binding proteins (RBPs), miRNAs, and related mRNAs to study and decipher the BC triggering biological processes and pathways.\n\nRooting from the identified total of 25 co-differentially expressed circRNAs (DECs) between triple negative (TN) and luminal A subtypes of BC from microarray analysis, five hub DECs (hsa_circ_0003227, hsa_circ_0001955, hsa_circ_0020080, hsa_circ_0001666, and hsa_circ_0065173) and top eleven RBPs (AGO1, AGO2, EIF4A3, FMRP, HuR (ELAVL1), IGF2BP1, IGF2BP2, IGF2BP3, EWSR1, FUS, and PTB) were explored to form the upper stream regulatory elements. All the hub circRNAs were regarded as super sponge having multiple miRNA response elements (MREs) for numerous miRNAs. Then four leading miRNAs (hsa-miR-149, hsa-miR-182, hsa-miR-383, and hsa-miR-873) accountable for BC progression were also introduced from merging several ceRNAs networks. The predicted 7- and 8-mer MREs matches between hub circRNAs and leading miRNAs ensured their enduring regulatory capability. The mined downstream mRNAs of the circRNAs-miRNAs network then were presented to STRING database to form the PPI network and deciphering the issue from another point of view. The BC interconnected enriched pathways and processes guarantee the merits of the ceRNAs networks members as targetable therapeutic elements.\n\nThis study suggested extensive panels of novel covering therapeutic targets that are in charge of BC progression in every aspect, hence their impressive role cannot be excluded and needs deeper empirical laboratory designs.

bioinformatics

Regulation of HMGB2 integrates ribosome biogenesis and innate immune responses to DNA

Ribosomes are universally important in biology and their production is dysregulated by developmental disorders, cancer, and virus infection. Although presumed required for protein synthesis, how ribosome biogenesis impacts virus reproduction and cell-intrinsic immune responses remains untested. Surprisingly, we find that restricting ribosome biogenesis stimulated human cytomegalovirus (HCMV) replication without suppressing translation. Interfering with ribosomal RNA (rRNA) accumulation triggered nucleolar stress and repressed expression of High Mobility Group Box 2 (HMGB2), a chromatin-associated protein that facilitates cytoplasmic double-stranded (ds) DNA-sensing by cGAS. Furthermore, it reduced cytoplasmic HMGB2 abundance and impaired induction of interferon beta (IFNB1) mRNA, which encodes a critical anti-proliferative, proinflammatory cytokine, in response to HCMV or dsDNA in uninfected cells. This establishes that rRNA accumulation regulates innate immune responses to dsDNA by controlling HMGB2 abundance. Moreover, it reveals that rRNA accumulation and/or nucleolar activity unexpectedly regulate dsDNA-sensing to restrict virus reproduction and regulate inflammation.

microbiology

Refine your search to explore more results.