Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Uncovering Robust Patterns of MicroRNA Co-Expression across Cancers using Bayesian Relevance Networks

Co-expression networks have long been used as a tool for investigating the molecular circuitry governing biological systems. However, most algorithms for constructing co-expression networks were developed in the microarray era, before high-throughput sequencing--with its unique statistical properties--became the norm for expression measurement. Here we develop Bayesian Relevance Networks, an algorithm that uses Bayesian reasoning about expression levels to account for the differing levels of uncertainty in expression measurements between highly- and lowly-expressed entities, and between samples with different sequencing depths. It combines data from groups of samples (e.g., replicates) to estimate group expression levels and confidence ranges. It then computes uncertainty-moderated estimates of cross-group correlations between entities, and uses permutation testing to assess their statistical significance. Using large scale miRNA data from The Cancer Genome Atlas, we show that our Bayesian update of the classical Relevance Networks algorithm provides improved reproducibility in co-expression estimates and lower false discovery rates in the resulting co-expression networks. Software is available at www.perkinslab.ca/Software.html.

bioinformatics

Protein Profiling In Cancer Cell Lines And Tumor Tissue Using Reverse Phase Protein Arrays

Reverse phase protein array (RPPA) technology is an antibody-based high-throughput assay for protein profiling of biological specimens that allows for many measurements with very small amounts of cell lysate. Here, we report the sensitivity, reproducibility, and accuracy of a particular RPPA platform called Zeptosens. We customized the RPPA protocol for our in-house setup, and measured more than 80 total protein and phospho-protein levels in various cancer samples, including cell lines, organoids, tumor chunks, core needle biopsies, and laser-capture microdissected tissue samples. We discuss pros and cons of the RPPA platform, and describe results from profiling 15 cancer cell line cells using RPPA.

systems biology

An open library of human kinase domain constructs for automated bacterial expression

Kinases play a critical role in many cellular signaling pathways and are dysregulated in a number of diseases, such as cancer, diabetes, and neurodegeneration. Since the FDA approval of imatinib in 2001, therapeutics targeting kinases now account for roughly 50% of current cancer drug discovery efforts. The ability to explore human kinase biochemistry, biophysics, and structural biology in the laboratory is essential to making rapid progress in understanding kinase regulation, designing selective inhibitors, and studying the emergence of drug resistance. While insect and mammalian expression systems are frequently used for the expression of human kinases, bacterial expression systems are superior in terms of simplicity and cost-effectiveness but have historically struggled with human kinase expression. Following the discovery that phosphatase coexpression could produce high yields of Src and Abl kinase domains in bacterial expression systems, we have generated a library of 52 His-tagged human kinase domain constructs that express above 2 {micro}g/mL culture in a simple automated bacterial expression system utilizing phosphatase coexpression (YopH for Tyr kinases, Lambda for Ser/Thr kinases). Here, we report a structural bioinformatics approach to identify kinase domain constructs previously expressed in bacteria likely to express well in a simple high-throughput protocol, experiments demonstrating our simple construct selection strategy selects constructs with good expression yields in a test of 84 potential kinase domain boundaries for Abl, and yields from a high-throughput expression screen of 96 human kinase constructs. Using a fluorescence-based thermostability assay and a fluorescent ATP-competitive inhibitor, we show that the highest-expressing kinases are folded and have well-formed ATP binding sites. We also demonstrate how the resulting expressing constructs can be used for the biophysical and biochemical study of clinical mutations by engineering a panel of 48 Src mutations and 46 Abl mutations via single-primer mutagenesis and screening the resulting library for expression yields. The wild-type kinase construct library is available publicly via Addgene, and should prove to be of high utility for experiments focused on drug discovery and the emergence of drug resistance.

Biochemistry

The innate immune DNA sensor cGAS is a negative regulator of DNA repair hence promotes genome instability and cell death

Stringent regulation of DNA repair is essential for organismal integrity, but the mechanisms are not fully understood. Cyclic cGMP-AMP synthase (cGAS), the DNA sensor that alerts the innate immune system to the presence of foreign or damaged self-DNA in the cytoplasm is critical for the outcome of infections, inflammatory diseases and cancer. Besides this cytoplasmic function as an innate immune sensor, whether cGAS fulfills other biological roles remains unknown. Here we report that cGAS has a distinct role in the nucleus: it inhibits homologous recombination DNA repair (HR) thereby promoting genome instability and associated micronuclear generation and mitotic death. We show that cGAS-mediated inhibition of HR requires its DNA binding and oligomerization but not its catalytic activity or the downstream innate immune signaling events. Mechanistically, we show that cGAS impede RAD51-mediated DNA strand invasion, a key step in HR. These results uncover a new function of cGAS relevant for understanding its involvement in genome instability- associated disorders.

cell biology

iOmicsPASS: a novel method for integration of multi-omics data over biological networks and discovery of predictive subnetworks

We developed iOmicsPASS, an intuitive method for network-based multi-omics data integration and detection of biological subnetworks for phenotype prediction. The method converts abundance measurements into co-expression scores of biological networks and uses a powerful phenotype prediction method adapted for network-wise analysis. Simulation studies show that the proposed data integration approach considerably improves the quality of predictions. We illustrate iOmicsPASS through the integration of quantitative multi-omics data using transcription factor regulatory network and protein-protein interaction network for cancer subtype prediction. Our analysis of breast cancer data identifies network signatures surrounding established markers of molecular subtypes. The analysis of colorectal cancer data highlights a protein interactome surrounding key proto-oncogenes as predictive features of subtypes, rendering them more biologically interpretable than the approaches integrating data without a priori relational information. However, the results indicate that current molecular subtyping is overly dependent on transcriptomic data and crude integrative analysis fails to account for molecular heterogeneity in other -omics data. The analysis also suggest that tumor subtypes are not mutually exclusive and future subtyping should therefore consider multiplicity in assignments.\n\nAvailability: https://github.com/cssblab/iOmicsPASS

systems biology

Multi-omic and multi-view clustering algorithms: review and cancer benchmark

High throughput experimental methods developed in recent years have been used to collect large biomedical omics datasets. Clustering of such datasets has proven invaluable for biological and medical research, and helped reveal structure in data from several domains. Such analysis is often based on investigation of a single omic. The decreasing cost and development of additional high throughput methods now enable measurement of multi-omic data. Clustering multi-omic data has the potential to reveal further systems-level insights, but raises computational and biological challenges. Here we review algorithms for multi-omics clustering, and discuss key issues in applying these algorithms. Our review covers methods developed specifically for multi-omic data as well as generic multi-view methods developed in the machine learning community for joint clustering of multiple data types.\n\nIn addition, using cancer data from TCGA, we perform an extensive benchmark spanning ten different cancer types, providing the first systematic benchmark comparison of leading multi-omics and multiview clustering algorithms. The results highlight several key questions regarding the use of single-vs. multi-omics, the choice of clustering strategy, the power of generic multi-view methods and the use of approximated p-values for gauging solution quality. Due to the rapidly increasing use of multi-omics data, these issues may be important for future progress in the field.

bioinformatics

Evolutionary forecasting of phenotypic and genetic outcomes of experimental evolution in Pseudomonas

Experimental evolution with microbes is often highly repeatable under identical conditions, suggesting the possibility to predict short-term evolution. However, it is not clear to what degree evolutionary forecasts can be extended to related species in non-identical environments, which would allow testing of general predictive models and fundamental biological assumptions. To develop an extended model system for evolutionary forecasting, we used previous data and models of the genotype-to-phenotype map from the wrinkly spreader system in Pseudomonas fluorescens SBW25 to make predictions of evolutionary outcomes on different biological levels for Pseudomonas protegens Pf-5. In addition to sequence divergence (78% amino acid and 81% nucleotide identity) for the genes targeted by mutations, these species also differ in the inability of Pf-5 to make cellulose, which is the main structural basis for the adaptive phenotype in SBW25. The experimental conditions were also changed compared to the SBW25 system to test the robustness of forecasts to environmental variation. Forty-three mutants with increased ability to colonize the air-liquid interface were isolated, and the majority had reduced motility and was partly dependent on the pel exopolysaccharide as a structural component. Most (38/43) mutations are expected to disrupt negative regulation of the same three diguanylate cyclases as in SBW25, with a smaller number of mutations in promoter regions, including that of an uncharacterized polysaccharide operon. A mathematical model developed for SBW25 predicted the order of the three main pathways and the genes targeted by mutations, but differences in fitness between mutants and mutational biases also appear to influence outcomes. Mutated regions in proteins could be predicted in most cases (16/22), but parallelism at the nucleotide level was low and mutational hot spots were not conserved. This study demonstrates the potential of short-term evolutionary forecasting in experimental populations and provides testable predictions for evolutionary outcomes in other Pseudomonas species. Author SummaryBiological evolution is often repeatable in the short-term suggesting the possibility of forecasting and controlling evolutionary outcomes. In addition to its fundamental importance for biology, evolutionary processes are at the core of several major societal problems, including infectious diseases, cancer and adaptation to climate change. Experimental evolution allows study of evolutionary processes in real time and seems like an ideal way to test the predictability of evolution and our ability to make forecasts. However, lack of model systems where forecasts can be extended to other species evolving under different conditions has prevented studies that first predict evolutionary outcomes followed by direct testing. We showed that a well-characterized bacterial experimental evolution system, based on biofilm formation by Pseudomonas fluorescens at the surface of static growth tubes, can be extended to the related species Pseudomonas protegens. We tested evolutionary forecasts experimentally and showed that mutations mainly appear in the predicted genes resulting in similar phenotypes. We also identified factors that we cannot yet predict, such as variation in mutation rates and differences in fitness. Finally, we make forecasts for other Pseudomonas species to be tested in future experiments.

evolutionary biology

Systems-Level Analysis Of 32 TCGA Cancers Reveals Disease-Dependent tRNA Fragmentation Patterns And Very Selective Associations With Messenger RNAs And Repeat Elements

We mined 10,274 datasets from The Cancer Genome Atlas (TCGA) for tRNA fragments (tRFs) that overlap nuclear and mitochondrial (MT) mature tRNAs. Across 32 cancer types, we identified 20,722 distinct tRFs, a third of which arise from MT tRNAs. Most of the fragments belong to the novel category of i-tRFs, i.e. they are wholly internal to the mature tRNAs. The abundances and cleavage patterns of the identified tRFs depend strongly on cancer type. Of note, in all 32 cancer types, we find that tRNAHisGTG produces multiple and abundant 5{acute}-tRFs with a uracil at the -1 position, instead of the expected post-transcriptionally-added guanosine. Strikingly, these -1U His 5{acute}tRFs are produced in ratios that remain constant across all analyzed normal and cancer samples, a property that makes tRNAHisGTG unique among all tRNAs. We also found numerous tRFs to be negatively correlated with many messenger RNAs (mRNAs) that belong primarily to four universal biological processes: transcription, cell adhesion, chromatin organization and development/morphogenesis. However, the identities of the mRNAs that belong to these processes and are negatively correlated with tRFs differ from cancer to cancer. Notably, the protein products of these mRNAs localize to specific cellular compartments, and do so in a cancer-dependent manner. Moreover, the genomic span of mRNAs that are negatively correlated with tRFs are enriched in multiple categories of repeat elements. Conversely, the genomic span of mRNAs that are positively correlated with tRFs are depleted in repeat elements. These findings suggest novel and far-reaching roles for tRFs and indicate their involvement in system-wide interconnections in the cell. All discovered tRFs from TCGA can be downloaded from https://cm.jefferson.edu/tcga-mintmap-profiles or studied interactively through the newly-designed version 2.0 of MINTbase at https://cm.jefferson.edu/MINTbase.\n\nNOTE: while the manuscript is under review, the content on the page https://cm.jefferson.edu/tcgamintmap-profiles is password protected and available only to Reviewers.\n\nKey PointsO_LIComplexity: tRNAs exhibit a complex fragmentation pattern into a multitude of tRFs that are conserved within the samples of a given cancer but differ across cancers.\nC_LIO_LIVery extensive mitochondrial contributions: the 22 tRNAs of the mitochondrion (MT) contribute 1/3rd of all tRFs found across cancers, a disproportionately high number compared to the tRFs from the 610 nuclear tRNAs.\nC_LIO_LIUridylated (not guanylated) 5{acute}-His tRFs: in all human tissues analyzed, tRNAHisGTG produces many abundant modified 5{acute}-tRFs with a U at their \"-1\" position (-1U 5{acute}-tRFs), instead of a G.\nC_LIO_LILikely central roles for tRNAHisGTG: the relative abundances of the -1U 5{acute}-tRFs from tRNAHisGTG remain strikingly conserved across the 32 cancers, a property that makes tRNAHisGTG unique among all tRNAs and isoacceptors.\nC_LIO_LISelective tRF-mRNA networks: tRFs are negatively correlated with mRNAs that differ characteristically from cancer to cancer.\nC_LIO_LIMitochondrion-encoded tRFs are associated with nuclear proteins: in nearly all cancers, and in a cancer-specific manner, tRFs produced by the 22 mitochondrial tRNAs are negatively correlated with mRNAs whose protein products localize to the nucleus.\nC_LIO_LItRFs are associated with membrane proteins: in all cancers, and in a cancer-specific manner, nucleus-encoded and MT-encoded tRFs are negatively correlated with mRNAs whose protein products localize to the cells membrane.\nC_LIO_LItRFs are associated with secreted proteins: in all cancers, and in a cancer-specific manner, nucleusencoded and MT-encoded tRFs are negatively correlated with mRNAs whose protein products are secreted from the cell.\nC_LIO_LItRFs are associated with numerous mRNAs through repeat elements: in all cancers, and in a cancerspecific manner, the genomic span of mRNAs that are negatively correlated with tRFs are enriched in specific categories of repeat elements.\nC_LIO_LIintra-cancer tRF networks can depend on sex and population origin: within a cancer, positive and negative tRF-tRF correlations can be modulated by patient attributes such as sex and population origin.\nC_LIO_LIweb-enabled exploration of an \"Atlas for tRFs\": we released a new version of MINTbase to provide users with the ability to study 26,531 tRFs compiled by mining 11,719 public datasets (TCGA and other sources).\nC_LI

systems biology

Inferring tree causal models of cancer progression with probability raising

Existing techniques to reconstruct tree models of progression for accumulative processes, such as cancer, seek to estimate causation by combining correlation and a frequentist notion of temporal priority. In this paper, we define a novel theoretical framework called CAPRESE (CAncer PRogression Extraction with Single Edges) to reconstruct such models based on the notion of probabilistic causation defined by Suppes. We consider a general reconstruction setting complicated by the presence of noise in the data due to biological variation, as well as experimental or measurement errors. To improve tolerance to noise we define and use a shrinkage-like estimator. We prove the correctness of our algorithm by showing asymptotic convergence to the correct tree under mild constraints on the level of noise. Moreover, on synthetic data, we show that our approach outperforms the state-of-the-art, that it is efficient even with a relatively small number of samples and that its performance quickly converges to its asymptote as the number of samples increases. For real cancer datasets obtained with different technologies, we highlight biologically significant differences in the progressions inferred with respect to other competing techniques and we also show how to validate conjectured biological relations with progression models.

Bioinformatics

Tensor decomposition--based unsupervised feature extraction for integrated analysis of TCGA data on microRNA expression and promoter methylation of genes in ovarian cancer

Integrated analysis of epigenetic profiles is important but difficult. Tensor decomposition-based unsupervised feature extraction was applied here to data on microRNA (miRNA) expression and promoter methylation of genes in ovarian cancer. It selected seven miRNAs and 241 genes by expression levels and promoter methylation degrees, respectively, such that they showed differences between eight normal ovarian tissue samples and 569 tumor samples. The expression levels of the seven miRNAs and the degrees of promoter methylation of the 241 genes also correlated significantly. Conventional Students t test-based feature selection failed to identify miRNAs and genes that have the above properties. On the other hand, biological evaluation of the seven identified miRNAs and 241 identified genes suggests that they are strongly related to cancer as expected.

bioinformatics

Population-level characterization of pathway alterations with SLAPenrich dissects heterogeneity of cancer hallmark acquisition

Cancer hallmarks are evolutionary traits required by a tumour to develop. While extensively characterised, the way these traits are achieved through the accumulation of somatic mutations in key biological pathways is not fully understood. To shed light on this subject, we characterised the landscape of pathway alterations associated with somatic mutations observed in 4,415 patients across ten cancer types, using 374 orthogonal pathway gene-sets mapped onto canonical cancer hallmarks. Towards this end, we developed SLAPenrich: a computational method based on population-level statistics, freely available as an open source R package. Assembling the identified pathway alterations into sets of hallmark signatures allowed us to connect somatic mutations to clinically interpretable cancer mechanisms. Further, we explored the heterogeneity of these signatures, in terms of ratio of altered pathways associated with each individual hallmark, assuming that this is reflective of the extent of selective advantage provided to the cancer type under consideration. Our analysis revealed the predominance of certain hallmarks in specific cancer types, thus suggesting different evolutionary trajectories across cancer lineages.\n\nFinally, although many pathway alteration enrichments are guided by somatic mutations in frequently altered high-confidence cancer genes, excluding these driver mutations preserves the hallmark heterogeneity signatures, thus the detected hallmarks predominance across cancer types. As a consequence, we propose the hallmark signatures as a ground truth to characterise tails of infrequent genomic alterations and identify potential novel cancer driver genes and networks.

Bioinformatics

INKA, an integrative data analysis pipeline for phosphoproteomic inference of active phosphokinases

Identifying (hyper)active kinases in cancer patient tumors is crucial to enable individualized treatment with specific inhibitors. Conceptually, kinase activity can be gleaned from global protein phosphorylation profiles obtained with mass spectrometry-based phosphoproteomics. A major challenge is to relate such profiles to specific kinases to identify (hyper)active kinases that may fuel growth/progression of individual tumors. Approaches have hitherto focused on phosphorylation of either kinases or their substrates. Here, we combine kinase-centric and substrate-centric information in an Integrative Inferred Kinase Activity (INKA) analysis. INKA utilizes label-free quantification of phosphopeptides derived from kinases, kinase activation loops, kinase substrates deduced from prior experimental knowledge, and kinase substrates predicted from sequence motifs, yielding a single score. This multipronged, stringent analysis enables ranking of kinase activity and visualization of kinase-substrate relation networks in a biological sample. As a proof of concept, INKA scoring of phosphoproteomic data for different oncogene-driven cancer cell lines inferred top activity of implicated driver kinases, and relevant quantitative changes upon perturbation. These analyses show the ability of INKA scoring to identify (hyper)active kinases, with potential clinical significance.

systems biology

Histone Deacetylase 11 is an ε-N-Myristoyllysine Hydrolase

Histone deacetylase (HDAC) enzymes are important regulators of diverse biological function, including gene expression, rendering them potential targets for intervention in a number of diseases, with a handful of compounds approved for treatment of certain hematologic cancers. Among the human zinc-dependent HDACs, the most recently discovered member, HDAC11, is the only member assigned to subclass IV, the smallest protein, and the least well understood with regards to biological function. Here we show that HDAC11 cleaves long chain acyl modifications on lysine side chains with remarkable efficiency compared to acetyl groups. We further show that several common types of HDAC inhibitors, including the approved drugs romidepsin and vorinostat, do not inhibit this enzymatic activity. Macrocyclic hydroxamic acid-containing peptides, on the other hand, potently inhibit HDAC11 demyristoylation activity. These findings should be taken carefully into consideration in future investigations of the biological function of HDAC11 and will serve as a foundation for the development of selective chemical probes targeting HDAC11.

biochemistry

MCbiclust: a novel algorithm to discover large-scale functionally related gene sets from massive transcriptomics data collections

The potential to understand fundamental biological processes from gene expression data has grown parallel with the recent explosion of the size of data collections. However, to exploit this potential, novel analytical methods are required, capable of handling massive data matrices. We found current methods limited in the size of correlated gene sets they could discover within biologically heterogeneous data collections, hampering the identification of multi-gene controlled fundamental cellular processes such as energy metabolism, organelle biogenesis and stress responses. Here we describe a novel biclustering algorithm called Massively Correlated Biclustering (MCbiclust) that selects samples and genes from large datasets with maximal correlated gene expression, allowing regulation of complex pathway to be examined. The method has been evaluated using synthetic data and applied to large bacterial and cancer cell datasets. We show that the large biclusters discovered, so far elusive to identification by existing techniques, are biologically relevant and thus MCbiclust has great potential use in the analysis of transcriptomics data to identify large scale unknown effects hidden within the data. The identified massive biclusters can be used to develop improved transcriptomics based diagnosis tools for diseases caused by altered gene expression, or used for further network analysis to understand genotype-phenotype correlations.

Bioinformatics

MINT: A multivariate integrative method to identify reproducible molecular signatures across independent experiments and platforms

BackgroundMolecular signatures identified from high-throughput transcriptomic studies often have poor reliability and fail to reproduce across studies. One solution is to combine independent studies into a single integrative analysis, additionally increasing sample size. However, the different protocols and technological platforms across transcriptomic studies produce unwanted systematic variation that strongly confounds the integrative analysis results. When studies aim to discriminate an outcome of interest, the common approach is a sequential two-step procedure; unwanted systematic variation removal techniques are applied prior to classification methods.\n\nResultsTo limit the risk of overfitting and over-optimistic results of a two-step procedure, we developed a novel multivariate integration method, MINT, that simultaneously accounts for unwanted systematic variation and identifies predictive gene signatures with greater reproducibility and accuracy. In two biological examples on the classification of three human cell types and four subtypes of breast cancer, we combined high-dimensional microarray and RNA-seq data sets and MINT identified highly reproducible and relevant gene signatures predictive of a given phenotype. MINT led to superior classification and prediction accuracy compared to the existing sequential two-step procedures.\n\nConclusionsMINT is a powerful approach and the first of its kind to solve the integrative classification framework in a single step by combining multiple independent studies. MINT is computationally fast as part of the mixOmics R CRAN package, available at http://www.mixOmics.org/mixMINT/ and http://cran.r-project.org/web/packages/mixOmics/.

Bioinformatics

Takeover times for a simple model of network infection

We study a stochastic model of infection spreading on a network. At each time step a node is chosen at random, along with one of its neighbors. If the node is infected and the neighbor is susceptible, the neighbor becomes infected. How many time steps T does it take to completely infect a network of N nodes, starting from a single infected node? An analogy to the classic \"coupon collector\" problem of probability theory reveals that the takeover time T is dominated by extremal behavior, either when there are only a few infected nodes near the start of the process or a few susceptible nodes near the end. We show that for N >> 1, the takeover time T is distributed as a Gumbel for the star graph; as the sum of two Gumbels for a complete graph and an Erd[o]s-Renyi random graph; as a normal for a one-dimensional ring and a two-dimensional lattice; and as a family of intermediate skewed distributions for d-dimensional lattices with d [≥] 3 (these distributions approach the sum of two Gumbels as d approaches infinity). Connections to evolutionary dynamics, cancer, incubation periods of infectious diseases, first-passage percolation, and other spreading phenomena in biology and physics are discussed.

evolutionary biology

Identification and characterization of moonlighting long non-coding RNAs based on RNA and protein interactome

Moonlighting proteins are a class of proteins having multiple distinct functions, which play essential roles in a variety of cellular and enzymatic functioning systems. Although there have long been calls for computational algorithms for the identification of moonlighting proteins, research on approaches to identify moonlighting long non-coding RNAs (lncRNAs) has never been undertaken. Here, we introduce a methodology, MoonFinder, for the identification of moonlighting lncRNAs. MoonFinder is a statistical algorithm identifying moonlighting lncRNAs without a priori knowledge through the integration of protein interactome, RNA-protein interactions, and functional annotation of proteins. We identify 155 moonlighting lncRNA candidates and uncover that they are a distinct class of lncRNAs characterized by specific sequence and cellular localization features. The non-coding genes that transcript moonlighting lncRNAs tend to have shorter but more exons and the moonlighting lncRNAs have a localization tendency of residing in the cytoplasmic compartment in comparison with the nuclear compartment. Moreover, moonlighting lncRNAs and moonlighting proteins are rather mutually exclusive in terms of both their direct interactions and interacting partners. Our results also shed light on how the moonlighting candidates and their interacting proteins implicated in the formation and development of cancers and other diseases.

systems biology

Distinct sequence patterns in the active postmortem transcriptome

Our previous study found more than 500 transcripts significantly increased in abundance in the zebrafish and mouse several hours to days postmortem relative to live controls. The current literature suggests that most mRNAs are post-transcriptionally regulated in stressful conditions, we rationalized that the postmortem transcripts must contain sequence features (3 to 9 mers) that are unique from those in the rest of the transcriptome - specifically, binding sites for proteins and/or non-coding RNAs involved in regulation. Our new study identified 5117 and 2245 over-represented sequence features in the mouse and zebrafish, respectively. Some of these features were disproportionately distributed along the transcripts with high densities in the 3-UTR region of the zebrafish (0.3 mers/nt) and the ORFs of the mouse (0.6 mers/nt). Yet, the highest density (2.3 mers/nt) occurred in the ORFs of 11 mouse transcripts that lacked UTRs. Our results suggest that these transcripts might serve as molecular sponges that sequester RNA binding proteins and/or microRNAs, increasing the stability and gene expression of other transcripts. In addition, some features were identified as binding sites for Rbfox and Hud proteins that are also involved in increasing transcript stability and gene expression. Hence, our results are consistent with the hypothesis that transcripts involved in responding to extreme stress have sequence features that make them different from the rest of the transcriptome, which presumably has implications for post-transcriptional regulation in disease, starvation, and cancer.\n\nABBREVIATIONS

systems biology