Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Population heterogeneity in the epithelial to mesenchymal transition is controlled by NFAT and phosphorylated Sp1

Epithelial to mesenchymal transition (EMT) is an essential differentiation program during tissue morphogenesis and remodeling. EMT is induced by soluble transforming growth factor {beta} (TGF-{beta}) family members, and restricted by vascular endothelial growth factor family members. While many downstream molecular regulators of EMT have been identified, these have been largely evaluated individually without considering potential crosstalk. In this study, we created an ensemble of dynamic mathematical models describing TGF-{beta} induced EMT to better understand the operational hierarchy of this complex molecular program. We used ordinary differential equations (ODEs) to describe the transcriptional and post-translational regulatory events driving EMT. Model parameters were estimated from multiple data sets using multiobjective optimization, in combination with cross-validation. TGF-{beta} exposure drove the model population toward a mesenchymal phenotype, while an epithelial phenotype was enhanced following vascular endothelial growth factor A (VEGF-A) exposure. Simulations predicted that the transcription factors phosphorylated SP1 and NFAT were master regulators promoting or inhibiting EMT, respectively. Surprisingly, simulations also predicted that a cellular population could exhibit phenotypic heterogeneity (characterized by a significant fraction of the population with both high epithelial and mesenchymal marker expression) if treated simultaneously with TGF-{beta} and VEGF-A. We tested this prediction experimentally in both MCF10A and DLD1 cells and found that upwards of 45% of the cellular population acquired this hybrid state in the presence of both TGF-{beta} and VEGF-A. We experimentally validated the predicted NFAT/Sp1 signaling axis for each phenotype response. Lastly, we found that cells in the hybrid state had significantly different functional behavior when compared to VEGF-A or TGF-{beta} treatment alone. Together, these results establish a predictive mechanistic model of EMT susceptibility, and potentially reveal a novel signaling axis which regulates carcinoma progression through an EMT versus tubulogenesis response.\n\nAuthor SummaryTissue formation and remodeling requires a complex and dynamic balance of interactions between epithelial cells, which reside on the surface, and mesenchymal cells that reside in the tissue interior. During embryonic development, wound healing, and cancer, epithelial cells transform into a mesenchymal cell to form new types of tissues. It is important to understand this process so that it can be controlled to generate beneficial effects and limit pathological differentiation. Much research over the past 20 years has identified many different molecular species that are relevant, but these have mainly been studied one at a time. In this study, we developed and implemented a novel computational strategy to interrogate the key players in this transformation process to identify which are the major bottlenecks. We determined that NFATc1 and pSP1 are essential for promoting epithelial or mesenchymal differentiation, respectively. We then predicted the existence of a partially transformed cell that exhibits both epithelial and mesenchymal characteristics. We found this partial cell type develops a network of invasive but stunted vascular structures that may be a unique cell target for understanding cancer progression and angiogenesis.

Systems Biology

Universal attenuators and their interactions with feedback loops in gene regulatory networks

Using a combination of mathematical modelling, statistical simulation and large-scale data analysis we study the properties of linear regulatory chains (LRCs) within gene regulatory networks (GRNs). Our modelling indicates that downstream genes embedded within LRCs are highly insulated from the variation in expression of upstream genes, and thus LRCs act as attenuators. This observation implies a progressively weaker functionality of LRCs as their length increases. When analysing the preponderance of LRCs in the GRNs of E. coli K12 and several other organisms, we find that very long LRCs are essentially absent. In both E. coli and M. tuberculosis we find that four-gene LRCs are intimately linked to identical feedback loops that are involved in potentially chaotic stress response, indicating that the dynamics of these potentially destabilising motifs are strongly restrained under homeostatic conditions. The same relationship is observed in a human cancer cell line (K562), and we postulate that four-gene LRCs act as \"universal attenuators\". These findings suggest a role for long LRCs in dampening variation in gene expression, thereby protecting cell identity, and in controlling dramatic shifts in cell-wide gene expression through inhibiting chaos-generating motifs.\n\nIn briefWe present a general principle that linear regulatory chains exponentially attenuate the range of expression in gene regulatory networks. The discovery of a universal interplay between linear regulatory chains and genetic feedback loops in microorganisms and a human cancer cell line is analysed and discussed.\n\nHighlightsWithin gene networks, linear regulatory chains act as exponentially strong attenuators of upstream variation\n\nBecause of their exponential behaviour, linear regulatory chains beyond a few genes provide no additional functionality and are rarely observed in gene networks across a range of different organisms\n\nNovel interactions between four-gene linear regulatory chains and feedback loops were discovered in E. coli, M. tuberculosis and human cancer cells, suggesting a universal mechanism of control.

Systems Biology

Stopping Transformed Growth with Cytoskeletal Proteins: Turning a Devil into an Angel

The major hallmark of cancer cells is uncontrollable growth on soft matrices (transformed growth), which indicates that they have lost the ability to properly sense the rigidity of their surroundings. Recent studies of fibroblasts show that local contractions by cytoskeletal rigidity sensor units block growth on soft surfaces and their depletion causes transformed growth. The contractile system involves many cytoskeletal proteins that must be correctly assembled for proper rigidity sensing. We tested the hypothesis that cancer cells lack rigidity sensing due to their inability to assemble contractile units because of altered cytoskeletal protein levels. In four widely different cancers, there were over ten-fold fewer rigidity-sensing contractions compared with normal fibroblasts. Restoring normal levels of cytoskeletal proteins restored rigidity sensing and rigidity-dependent growth in transformed cells. Most commonly, this involved restoring balanced levels of the tropomyosins 2.1 (often depleted by miR-21) and 3 (often overexpressed). Restored cells could be transformed again by depleting other cytoskeletal proteins including myosin IIA. Thus, the depletion of rigidity sensing modules enables growth on soft surfaces and many different perturbations of cytoskeletal proteins can disrupt rigidity sensing thereby causing transformed growth of cancer cells.

cell biology

Golgi Renaissance: the pivotal role of the largest Golgi protein giantin

Golgi undergoes disorganization in response to the drugs or alcohol, but it is able to restore compact structure under recovery. This self-organization mechanism remains mostly elusive, as does the role of giantin, the largest Golgi matrix dimeric protein. Here, we found that in cells treated with Brefeldin A (BFA) or ethanol (EtOH), Golgi disassembly is associated with giantin de-dimerization, which was restored to the dimer form after BFA or EtOH washout. Cells lacking giantin are disabled for the restoration of the classical ribbon Golgi, and they demonstrate altered trafficking of proteins to the cell surface. The fusion of the nascent Golgi membranes is mediated by the cross-membrane interaction of Rab6a GTPase and giantin. Giantin is involved in the formation of long intercisternal connections, which in giantin-depleted cells was replaced by the short bridges that formed via oligomerization of GRASP65. This phenomenon occurs in advanced prostate cancer cells, in which a fragmented Golgi phenotype is maintained by the dimerization of GRASP65. Thus, we provide a model of Golgi Renaissance, which is impaired in aggressive prostate cancer.

cell biology

Delivery of GalNAc-conjugated splice-switching ASOs to non-hepatic cells through ectopic expression of asialoglycoprotein receptor

Splice-switching antisense oligonucleotides (ASOs) are promising therapeutic tools to target various genetic diseases, including cancer. However, in vivo delivery of ASOs to orthotopic tumors in cancer mouse models or to certain target tissues remains challenging. A viable solution already in use is receptor-mediated uptake of ASOs via tissue-specific receptors. For example, the asialoglycoprotein receptor (ASGP-R) is exclusively expressed in hepatocytes. Triantennary GalNAc (GN3)-conjugated ASOs bind to the receptor and are efficiently internalized by endocytosis, enhancing ASO potency in the liver. Here we explore the use of GalNAc-mediated targeting to deliver therapeutic splice-switching ASOs to cancer cells that ectopically express ASGP-R, both in vitro and in tumor mouse models. We found that ectopic expression of the major isoform ASGP-R1 H1a is sufficient to promote uptake and increase GN3-ASO potency to various degrees in all tested cancer cells. We show that cell-type specific glycosylation of the receptor does not affect its activity. In vivo, GN3-conjugated ASOs specifically target subcutaneous xenograft tumors that ectopically express ASGP-R1, and modulate splicing significantly more strongly than unconjugated ASOs. Our work shows that GN3-targeting is a useful tool for proof-of-principle studies in orthotopic cancer models, until endogenous receptors are identified and exploited for efficiently targeting cancer cells.

molecular biology

Comprehensive functional profiling of long non-coding RNAs through a novel pan-cancer integration approach and modular analysis of their protein-coding gene association networks

Long non-coding RNAs (lncRNAs) are emerging as crucial regulators of cellular processes in diseases such as cancer, although the functions of most remain poorly understood. To address this, here we apply a novel strategy to integrate gene expression profiles across 32 cancer types, and cluster human lncRNAs based on their pan-cancer protein-coding gene associations. By doing so, we derive 16 lncRNA modules whose unique properties allow simultaneous inference of function, disease specificity and regulation for over 800 lncRNAs. Remarkably, modules could be grouped into just four functional themes: transcription regulation, immunological, extracellular, and neurological, with module generation frequently driven by lncRNA tissue specificity. Notably, three modules associated with the extracellular matrix represented potential networks of lncRNAs regulating key events in tumour progression. These included a tumour-specific signature of 33 lncRNAs that may play a role in inducing epithelialmesenchymal transition through modulation of TGF{beta} signalling, and two stromal-specific modules comprising 26 lncRNAs linked to a tumour suppressive microenvironment, and 12 lncRNAs related to cancer-associated fibroblasts. At least one member of the 12-lncRNA signature was experimentally supported by siRNA knockdown, which resulted in attenuated differentiation of quiescent fibroblasts to a cancer-associated phenotype. Overall, the study provides a unique pan-cancer perspective on the lncRNA functional landscape, acting as a global source of novel hypotheses on lncRNA contribution to tumour progression.\n\nAuthor SummaryThe established view of protein production is that genomic DNA is transcribed into RNA, which is then translated into protein. Proteins play a critical role in shaping the function of each individual cell in the human body yet they represent less than 2% of human genomic sequence whilst up to 90% of the genome is transcribed. To explain this disparity, the existence of thousands of long non-coding RNAs (lncRNAs) has emerged that do not encode proteins but perform function as an RNA molecule. Most lncRNAs have yet to be assigned a specific biological role, so to address this we apply a novel computational approach to characterise the function of >800 lncRNAs through consistent association with protein coding genes across multiple cancer types. By doing so, we discover 16 \"modules\" of closely related lncRNAs that share broad functional themes, the most compelling of which consists of 12 lncRNAs that could regulate activation of specific cells neighbouring the tumour, leading to accelerated tumour progression and invasion. Overall, the study provides the most robust view of the lncRNA-protein coding gene landscape to date, adding to growing evidence that lncRNAs are key regulators of cancer, and have therapeutic potential comparable to proteins.

bioinformatics

Creating Standards for Evaluating Tumour Subclonal Reconstruction

Tumours evolve through time and space. Computational techniques have been developed to infer their evolutionary dynamics from DNA sequencing data. A growing number of studies have used these approaches to link molecular cancer evolution to clinical progression and response to therapy. There has not yet been a systematic evaluation of methods for reconstructing tumour subclonality, in part due to the underlying mathematical and biological complexity and to difficulties in creating gold-standards. To fill this gap, we systematically elucidated the key algorithmic problems in subclonal reconstruction and developed mathematically valid quantitative metrics for evaluating them. We then created approaches to simulate realistic tumour genomes, harbouring all known mutation types and processes both clonally and subclonally. We then simulated 580 tumour genomes for reconstruction, varying tumour read-depth and benchmarking somatic variant detection and subclonal reconstruction strategies. The inference of tumour phylogenies is rapidly becoming standard practice in cancer genome analysis; this study creates a baseline for its evaluation.

bioinformatics

Comparing alternative pipelines for cross-platform microarray gene expression data integration with RNA-seq data in breast cancer

BackgroundAccording to major public repositories statistics an overwhelming majority of the existing and newly uploaded data originates from microarray experiments. Unfortunately, the potential of this data to bring new insights is limited by the effects of individual study-specific biases due to small number of biological samples. Increasing sample size by direct microarray data integration increases the statistical power to obtain a more precise estimate of gene expression in a population of individuals resulting in lower false discovery rates. However, despite numerous recommendations for gene expression data integration, there is a lack of a systematic comparison of different processing approaches aimed to asses microarray platforms diversity and ambiguous probesets to genes correspondence, leading to low number of studies applying integration.\n\nResultsHere, we investigated five different approaches of the microarrays data processing in comparison with RNA-seq data on breast cancer samples. We aimed to evaluate different probesets annotations as well as different procedures of choosing between probesets mapped to the same gene. We show that pipelines rankings are mostly preserved across Affymetrix and Illumina platforms. BrainArray approach based on updated annotation and redesigned probesets definition and choosing probeset with the maximum average signal across the samples have best correlation with RNA-seq, while averaging probesets signals as well as scoring the quality of probes sequences mapping to the transcripts of the targeted gene have worse correlation. Finally, randomly selecting probeset among probesets mapped to the same gene significantly decreases the correlation with RNA-seq.\n\nConclusionWe show that methods, which rely on actual probesets signal intensities, are advantageous to methods considering biological characteristics of the probes sequences only and that cross-platform integration of datasets improves correlation with the RNA-seq data. We consider the results obtained in this paper contributive to the integrative analysis as a worthwhile alternative to the classical meta-analysis of the multiple gene expression datasets.

Bioinformatics

Independent component analysis provides clinically relevant insights into the biology of melanoma patients

The integration of publicly available and new patient-derived transcriptomic datasets is not straightforward and requires specialized approaches to deal with heterogeneity at technical and biological levels. Here we present a methodology that can overcome technical biases, predict clinically relevant outcomes and identify tumour-related biological processes in patients using previously collected large reference datasets. The approach is based on independent component analysis (ICA) - an unsupervised method of signal deconvolution. We developed parallel consensus ICA that robustly decomposes merged new and reference datasets into signals with minimal mutual dependency. By applying the method to a small cohort of primary melanoma and control samples combined with a large public melanoma dataset, we demonstrate that our method distinguishes cell-type specific signals from technical biases and allows to predict clinically relevant patient characteristics. Cancer subtypes, patient survival and activity of key tumour-related processes such as immune response, angiogenesis and cell proliferation were characterized. Additionally, through integration of transcriptomes and miRNomes, the method identified biological functions of miRNAs, which would otherwise not be possible.

genomics

EMT network-based feature selection improves prognosis prediction in lung adenocarcinoma

Various feature selection algorithms have been proposed to identify cancer prognostic biomarkers. In recent years, however, their reproducibility is criticized. The performance of feature selection algorithms is shown to be affected by the datasets, underlying networks and evaluation metrics. One of the causes is the curse of dimensionality, which makes it hard to select the features that generalize well on independent data. Even the integration of biological networks does not mitigate this issue because the networks are large and many of their components are not relevant for the phenotype of interest. With the availability of multi-omics data, integrative approaches are being developed to build more robust predictive models. In this scenario, the higher data dimensions create greater challenges.\n\nWe proposed a phenotype relevant network-based feature selection (PRNFS) framework and demonstrated its advantages in lung cancer prognosis prediction. We constructed cancer prognosis relevant networks based on epithelial mesenchymal transition (EMT) and integrated them with different types of omics data for feature selection. With less than 2.5% of the total dimensionality, we obtained EMT prognostic signatures that achieved remarkable prediction performance (average AUC values >0.8), very significant sample stratifications, and meaningful biological interpretations. In addition to finding EMT signatures from different omics data levels, we combined these single-omics signatures into multi-omics signatures, which improved sample stratifications significantly. Both single- and multi-omics EMT signatures were tested on independent multi-omics lung cancer datasets and significant sample stratifications were obtained.

bioinformatics

Predicting Carriers of Ongoing Selective Sweeps Without Knowledge of the Favored Allele

Methods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying many selective sweeps. For most of these sweeps, the favored allele remains unknown, making it difficult to distinguish carriers of the sweep from non-carriers. In an ongoing selective sweep, carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones. The main contribution of this paper is the development and analysis of a new statistic, the Haplotype Allele Frequency (HAF) score. The HAF score, assigned to individual haplotypes in a sample, naturally captures many of the properties shared by haplotypes carrying a favored allele. We provide a theoretical framework for computing expected HAF scores under different evolutionary scenarios, and we validate the theoretical predictions with simulations. As an application of HAF score computations, we develop an algorithm (PreCIOSS: Predicting Carriers of Ongoing Selective Sweeps) to identify carriers of the favored allele in selective sweeps, and we demonstrate its power on simulations of both hard and soft sweeps, as well as on data from well-known sweeps in human populations.\n\nAuthor summaryMethods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying genomic regions under positive selection. However, methods that detect positive selective sweeps do not typically identify the favored allele, or even the haplotypes carrying the favored allele. The main contribution of this paper is the development and analysis of a new statistic (the HAF score), assigned to individual haplotypes. Using both theoretical analyses and simulations, we describe how the HAF scores differ for carriers and non-carriers of the favored allele, and how they change dynamically during a selective sweep. We also develop an algorithm, PreCIOSS, for separating carriers and non-carriers. Our tool has broad applicability as carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones.

Evolutionary Biology

CD44 Controls Endothelial Proliferation and Functions as Endogenous Inhibitor of Angiogenesis

CD44 transmembrane glycoprotein is involved in angiogenesis, but it is not clear whether CD44 functions as a pro- or antiangiogenic molecule. Here, we assess the role of CD44 in angiogenesis and endothelial proliferation by using Cd44-null mice and CD44 silencing in human endothelial cells. We demonstrate that angiogenesis is increased in Cd44-null mice compared to either wild-type or heterozygous animals. Silencing of CD44 expression in cultured endothelial cells results in their augmented proliferation and viability. The growth-suppressive effect of CD44 is mediated by its extracellular domain and is independent of its hyaluronan binding function. CD44-mediated effect on cell proliferation is independent of specific angiogenic growth factor stimulation. These results show that CD44 expression on endothelial cells constrains endothelial cell proliferation and angiogenesis. Thus, endothelial CD44 might serve as a therapeutic target both in the treatment of cardiovascular diseases, where endothelial protection is desired, as well as in cancer treatment, due to its antiangiogenic properties.

Cell Biology

QuickRNASeq: Guide For Pipeline Implementation And For Interactive Results Visualization

i.Summary/AbstractSequencing of transcribed RNA molecules (RNA-seq) has been used wildly for studying cell transcriptomes in bulk or at the single-cell level (1, 2, 3) and is becoming the de facto technology for investigating gene expression level changes in various biological conditions, on the time course, and under drug treatments. Furthermore, RNA-Seq data helped identify fusion genes that are related to certain cancers (4). Differential gene expression before and after drug treatments provides insights to mechanism of action, pharmacodynamics of the drugs, and safety concerns (5). Because each RNA-seq run generates tens to hundreds of millions of short reads with size ranging from 50bp-200bp, a tool that deciphers these short reads to an integrated and digestible analysis report is in high demand. QuickRNASeq (6) is an application for large-scale RNA-seq data analysis and real-time interactive visualization of complex data sets. This application automates the use of several of the best open-source tools to efficiently generate user friendly, easy to share, and ready to publish report. Figure 1 illustrates some of the interactive plots produced by QuickRNASeq. The visualization features of the application have been further improved since its first publication in early 2016. The original QuickRNASeq publication (6) provided details of background, software selection, and implementation. Here, we outline the steps required to implement QuickRNASeq in users own environment, as well as demonstrate some basic yet powerful utilities of the advanced interactive visualization modules in the report.\n\nO_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC=\"FIGDIR/small/125856_fig1.gif\" ALT=\"Figure 1\">\nView larger version (59K):\norg.highwire.dtl.DTLVardef@1f6fb70org.highwire.dtl.DTLVardef@1f5a748org.highwire.dtl.DTLVardef@b990fborg.highwire.dtl.DTLVardef@dd5336_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 1C_FLOATNO Interactive plots from QuickRNASeq report. Figures (a, b, c) can be retrieved by clicking on the pointing hands as shown in figure Id. On any of these interactive plots, mouse over each sample displays associated sample QC metrics, (a) Read mapping summary in the expanded display mode, (b) SNP concordance matrix of 48 samples from 5 donors. Samples from the same donor should be highly concordant, (c) Gene expression chart, which shows the number of genes past various expression thresholds, (d) Center portion of the QuickRNASeq report, (e) Parallel plot linking multiple QC measures for the same samples plus table of multi-dimensional QC measures.\n\nC_FIG

bioinformatics

LIN28 selectively modulates a subclass of let-7 microRNAs

LIN28 is a bipartite RNA-binding protein that post-transcriptionally inhibits let-7 microRNAs to regulate development and influence disease states. However, the mechanisms of let-7 suppression remains poorly understood, because LIN28 recognition depends on coordinated targeting by both the zinc knuckle domain (ZKD)--which binds a GGAG-like element in the precursor--and the cold shock domain (CSD), whose binding sites have not been systematically characterized. By leveraging single-nucleotide-resolution mapping of LIN28 binding sites in vivo, we determined that the CSD recognizes a (U)GAU motif. This motif partitions the let-7 family into Class I precursors with both CSD and ZKD binding sites and Class II precursors with ZKD but no CSD binding sites. LIN28 in vivo recognition--and subsequent 3' uridylation and degradation--of Class I precursors is more efficient, leading to their stronger suppression in LIN28-activated cells and cancers. Thus, CSD binding sites amplify the effects of the LIN28 activation with potential implication in development and cancer.

molecular biology

Genomic copy-number loss is rescued by self-limiting production of DNA circles

Copy-number changes generate phenotypic variability in health and disease. Whether organisms protect against copy-number changes is largely unknown. Here, we show that Saccharomyces cerevisiae monitors the copy number of its ribosomal DNA (rDNA) and rapidly responds to copy-number loss with the clonal amplification of extrachromosomal rDNA circles (ERCs) from chromosomal repeats. ERC production is proportional to repeat loss and reaches a dynamic steady state that responds to the addition of exogenous rDNA copies. ERC levels are also modulated by RNAPI activity and diet, suggesting that rDNA copy number is calibrated against the cellular demand for rRNA. Lastly, we show that ERCs reinsert into the genome in a dosage-dependent manner, indicating that they provide a reservoir for ultimately increasing rDNA array length. Our results reveal a DNA-based mechanism for rapidly restoring copy number in response to catastrophic gene loss that shares fundamental features with unscheduled copy-number amplifications in cancer cells.

molecular biology

cis-regulatory architecture of a short-range EGFR organizing center in the Drosophila melanogaster leg.

We characterized the establishment of an Epidermal Growth Factor Receptor (EGFR) organizing center (EOC) during leg development in Drosophila melanogaster. Initial EGFR activation occurs in the center of leg discs by expression of the EGFR ligand Vn and the EGFR ligand-processing protease Rho, each through single enhancers, vnE and rhoE, that integrate inputs from Wg, Dpp, Dll and Sp1. Deletion of vnE and rhoE eliminates vn and rho expression in the center of the leg imaginal discs, respectively. Animals with deletions of both vnE and rhoE (but not individually) show distal but not medial leg truncations, suggesting that the distal source of EGFR ligands acts at short-range to only specify distal-most fates, and that multiple additional ring enhancers are responsible for medial fates. Further, based on the cis-regulatory logic of vnE and rhoE we identified many additional leg enhancers, suggesting that this logic is broadly used by many genes during Drosophila limb development.\n\nAuthor SummaryThe EGFR signaling pathway plays a major role in innumerable developmental processes in all animals and its deregulation leads to different types of cancer, as well as many other developmental diseases in humans. Here we explored the integration of inputs from the Wnt- and TGF-beta signaling pathways and the leg-specifying transcription factors Distal-less and Sp1 at enhancer elements of EGFR ligands. These enhancers trigger a specific EGFR-dependent developmental output in the fly leg that is limited to specifying distal-most fates. Our findings suggest that activation of the EGFR pathway during fly leg development occurs through the activation of multiple EGFR ligand enhancers that are active at different positions along the proximo-distal axis. Similar enhancer elements are likely to control EGFR activation in humans as well. Such DNA elements might be hot spots that cause formation of EGFR-dependent tumors if mutations in them occur. Thus, understanding the molecular characteristics of such DNA elements could facilitate the detection and treatment of cancer.

developmental biology

Learning common and specific patterns from data of multiple interrelated biological scenarios with matrix factorization

High-throughput biological technologies (e.g., ChIP-seq, RNA-seq and single-cell RNA-seq) rapidly accelerate the accumulation of genome-wide omics data in diverse interrelated biological scenarios (e.g., cells, tissues and conditions). Data dimension reduction and differential analysis are two common paradigms for exploring and analyzing such data. However, they are typically used in a separate or/and sequential manner. In this study, we propose a flexible non-negative matrix factorization framework CSMF to combine them into one paradigm to simultaneously reveal common and specific patterns from data generated under interrelated biological scenarios. We demonstrate the effectiveness of CSMF with four applications including pairwise ChIP-seq data describing the chromatin modification map on protein-DNA interactions between K562 and Huvec cell lines; pairwise RNA-seq data representing the expression profiles of two cancers (breast invasive carcinoma and uterine corpus endometrial carcinoma); RNA-seq data of three breast cancer subtypes; and single-cell sequencing data of human embryonic stem cells and differentiated cells at six time points. Extensive analysis yields novel insights into hidden combinatorial patterns embedded in these interrelated multi-modal data. Results demonstrate that CSMF is a powerful tool to uncover common and specific patterns with significant biological implications from data of interrelated biological scenarios.

bioinformatics

Inferring propensity amongst lung and breast carcinomas via overlapped gene expression profiles

Reconstruction of biological networks for topological analyses helps in correlation identification between various types of biomarkers. These networks have been vital components of System Biology in present era. Genes are the basic physical and structural unit of heredity. Genes act as instructions to make molecules called proteins. Alterations in the normal sequence of these genes are the root cause of various diseases and cancer is the prominent example disease caused by gene alteration or mutation. These slight alterations can be detected by microarray analysis. The high throughput data obtained by microarray experiments aid scientists in reconstructing cancer specific gene regulatory networks. The purpose of experiment performed is to find out the overlapping of the gene expression profiles of breast and lung cancer data, so that the common hub genes can be sifted and utilized as drug targets which could be used for the treatment of diseased conditions. In this study, first the differentially expressed genes have been identified (lung cancer and breast cancer), followed by a filtration approach and most significant genes are chosen using paired t-test and gene regulatory network construction. The obtained result has been checked and validated with the available databases and literature.

systems biology