Search bioRxivSearch

Biology subjects

Bouhaddou, M.

Publications and source records attributed to Bouhaddou, M..

6 recordsLinked to original sources

Gene-Specific Predictability of Protein Levels from mRNA Data in Humans

Transcriptomic data are widely available, and the extent to which they are predictive of protein abundances remains debated. Using multiple public databases, we calculate mRNA and mRNA-to-protein ratio variability across human tissues to quantify and classify genes for protein abundance predictability confidence. We propose that such predictability is best understood as a spectrum. A gene-specific, tissue-independent mRNA-to-protein ratio plus mRNA levels explains [~]80% of protein abundance variance for more predictable genes, as compared to [~]55% for less predictable genes. Protein abundance predictability is consistent with independent mRNA and protein data from two disparate cell lines, and mRNA-to-protein ratios estimated from publicly-available databases have predictive power in these independent datasets. Genes with higher predictability are enriched for metabolic function, tissue development/cell differentiation roles, and transmembrane transporter activity. Genes with lower predictability are associated with cell adhesion, motility and organization, the immune system, and the cytoskeleton. Surprisingly, many genes that regulate mRNA-to-protein ratios are constitutively expressed but also exhibit ratio variability, suggesting a general autoregulation mechanism whereby protein expression profile changes can be implemented quickly, or homeostatic sensing stabilizes protein abundances under fluctuating conditions. Gene classifications and their mRNA-to-protein ratios are provided as a resource to facilitate protein abundance predictions by others.

systems biology

Fluorescence Multiplexing with Spectral Imaging and Combinatorics

Ultraviolet-to-infrared fluorescence is a versatile and accessible assay modality, but is notoriously hard to multiplex due to overlap of wide emission spectra. We present an approach for fluorescence multiplexing using spectral imaging and combinatorics (MuSIC). MuSIC consists of creating new independent probes from covalently-linked combinations of individual fluorophores, leveraging the wide palette of currently available probes with the mathematical power of combinatorics. Probe levels in a mixture can be inferred from spectral emission scanning data. Theory and simulations suggest MuSIC can increase fluorescence multiplexing ~4-5 fold using currently available dyes and measurement tools. Experimental proof-of-principle demonstrates robust demultiplexing of nine solution-based probes using ~25% of the available excitation wavelength window (380-480 nm), consistent with theory. The increasing prevalence of white lasers, angle filter-based wavelength scanning, and large, sensitive multi-anode photo-multiplier tubes make acquisition of such MuSIC-compatible datasets increasingly attainable.

biochemistry

Network Reconstruction from Perturbation Time Course Data

Networks underlie much of biology from subcellular to ecological scales. Yet, understanding what experimental data are needed and how to use them for unambiguously identifying the structure of even small networks remains a broad challenge. Here, we integrate a dynamic least squares framework into established modular response analysis (DL-MRA), that specifies sufficient experimental perturbation time course data to robustly infer arbitrary two and three node networks. DL-MRA considers important network properties that current methods often struggle to capture: (i) edge sign and directionality; (ii) cycles with feedback or feedforward loops including self-regulation; (iii) dynamic network behavior; (iv) edges external to the network; and (v) robust performance with experimental noise. We evaluate the performance of and the extent to which the approach applies to cell state transition networks, intracellular signaling networks, and gene regulatory networks. Although signaling networks are often an application of network reconstruction methods, the results suggest that only under quite restricted conditions can they be robustly inferred. For gene regulatory networks, the results suggest that incomplete knockdown is often more informative than full knockout perturbation, which may change experimental strategies for gene regulatory network reconstruction. Overall, the results give a rational basis to experimental data requirements for network reconstruction and can be applied to any such problem where perturbation time course experiments are possible.

systems biology

Analysis of Copy Number Loss of the ErbB4 Receptor Tyrosine Kinase in Glioblastoma

Current treatments for glioblastoma multiforme (GBM)--an aggressive form of brain cancer--are minimally effective and yield a median survival of 14.6 months and a two-year survival rate of 30%. Given the severity of GBM and the limitations of its treatment, there is a need for the discovery of novel drug targets for GBM and more personalized treatment approaches based on the characteristics of an individuals tumor. Most receptor tyrosine kinases--such as EGFR--act as oncogenes, but publicly available data from the Cancer Cell Line Encyclopedia (CCLE) indicates copy number loss in the ERBB4 RTK gene across dozens of GBM cell lines, suggesting a potential tumor suppressor role. This loss is mutually exclusive with loss of its cognate ligand NRG1 in CCLE as well, more strongly suggesting a functional role. The availability of higher resolution copy number data from clinical GBM patients in The Cancer Genome Atlas (TCGA) revealed that a region in Intron 1 of the ERBB4 gene was deleted in 69.1% of tumor samples harboring ERBB4 copy number loss; however, it was also found to be deleted in the matched normal tissue samples from these GBM patients (n = 81). Using the DECIPHER Genome Browser, we also discovered that this mutation occurs at approximately the same frequency in the general population as it does in the disease population. We conclude from these results that this loss in Intron 1 of the ERBB4 gene is neither a de novo driver mutation nor a predisposing factor to GBM, despite the indications from CCLE. A biological role of this significantly occurring genetic alteration is still unknown. While this is a negative result, the broader conclusion is that while copy number data from large cell line-based data repositories may yield compelling hypotheses, careful follow up with higher resolution copy number assays, patient data, and general population analyses are essential to codify initial hypotheses.

systems biology

An Integrated Mechanistic Model of Pan-Cancer Driver Pathways Predicts Stochastic Proliferation and Death

Most cancer cells harbor multiple drivers whose epistasis and interactions with expression context clouds drug sensitivity prediction. We constructed a mechanistic computational model that is context-tailored by omics data to capture regulation of stochastic proliferation and death by pan-cancer driver pathways. Simulations and experiments explore how the coordinated dynamics of RAF/MEK/ERK and PI-3K/AKT kinase activities in response to synergistic mitogen or drug combinations control cell fate in a specific cellular context. In this context, synergistic ERK and AKT inhibitor-induced death is likely mediated by BIM rather than BAD. AKT dynamics explain S-phase entry synergy between EGF and insulin, but stochastic ERK dynamics seem to drive cell-to-cell proliferation variability, which in simulations are predictable from pre-stimulus fluctuations in C-Raf/B-Raf levels. Simulations predict MEK alteration negligibly influences transformation, consistent with clinical data. Our model mechanistically interprets context-specific landscapes between driver pathways and cell fates, moving towards more rational cancer combination therapy.

systems biology

A Comparison of mRNA Sequencing with Random Primed and 3’-Directed Libraries

Deep mRNA sequencing (mRNAseq) is the state-of-the-art for whole transcriptome measurements. A key step is creating a library of cDNA sequencing fragments from RNA. This is generally done by random priming, creating multiple sequencing fragments along the length of each transcript. A 3 end-focused library approach cannot detect differential splicing, but has potentially higher throughput at lower cost (~10-fold lower), along with the ability to improve quantification by using transcript molecule counting with unique molecular identifiers (UMI) to correct for PCR bias. Here, we compare implementation of such a 3-digital gene expression (3-DGE) approach with \"conventional\" random primed mRNAseq, which has not yet been done. We find that while conventional mRNAseq detects ~15% more genes, the resulting lists of differentially expressed genes and therefore biological conclusions and gene signatures are highly concordant between the two techniques. We also find good quantitative agreement on the level of individual genes between the two techniques in terms of both read counts and fold change between two conditions. We conclude that for high-throughput applications, the potential cost savings associated with the 3-DGE approach are a very reasonable tradeoff for modest reduction in sensitivity and inability to observe alternative splicing, and should enable much larger scale studies focused on not only differential expression analysis, but also quantitative transcriptome profiling. The computational scripts and programs, along with experimental standard operating procedures used in our pipeline presented here, are freely available on our website (www.dtoxs.org).

genomics