Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

Phenotype prediction in an Escherichia coli strain panel

Understanding how genetic variation contributes to phenotypic differences is a fundamental question in biology. Here, we set to predict fitness defects of an individual using mechanistic models of the impact of genetic variants combined with prior knowledge of gene function. We assembled a diverse panel of 696 Escherichia coli strains for which we obtained genomes and measured growth phenotypes in 214 conditions. We integrated variant effect predictors to derive gene-level probabilities of loss of function for every gene across strains. We combined these probabilities with information on conditional gene essentiality in the reference K-12 strain to predict the strains growth defects, providing significant predictions for up to 38% of tested conditions. The putative causal variants were validated in complementation assays highlighting commonly perturbed pathways in evolution for the emergence of growth phenotypes. Altogether, our work illustrates the power of integrating high-throughput gene function assays to predict the phenotypes of individuals.\n\nHighlightsO_LIAssembled a reference panel of E. coli strains\nC_LIO_LIGenotyped and high-throughput phenotyped the E. coli reference strain panel\nC_LIO_LIReliably predicted the impact of genetic variants in up to 38% of tested conditions\nC_LIO_LIHighlighted common genetic pathways for the emergence of deleterious phenotypes\nC_LI

systems biology

Methane Reduction Potential of Two Pacific Coast Macroalgae During in-vitro Ruminant Fermentation.

With increasing interest in feed based methane mitigation strategies, fueled by local legal directives aimed at methane production from the agricultural sector in California, identifying local sources of biological feed additives will be critical in keeping the implementation of these strategies affordable. In a recent study, the red alga Asparagopsis taxiformis stood out as the most effective species of seaweed to reduce methane production from enteric fermentation. Due to the potential differences in effectiveness based on the location from where A. taxiformis is collected and the financial burden of collection and transport, we tested the potential of A. taxiformis, as well as the brown seaweed Zonaria farlowii collected in the nearshore waters off Santa Catalina Island, CA, USA, for their ability to mitigate methane production during in-vitro rumen fermentation. At a dose rate of 5% dry matter (DM), A. taxiformis reduced methane production by 74% (p [≤] 0.01) and Z. farlowii reduced methane production by 11% (p [≤] 0.04) after 48 hours and 24 hours of in-vitro rumen fermentation respectively. The methane reducing effect of A. taxiformis and Z. farlowii described here make these local macroalgae promising candidates for biotic methane mitigation strategies in the largest milk producing state in the US. To determine their real potential as methane mitigating feed supplements in the dairy industry, their effect in-vivo requires investigation.

systems biology

Integration of single-cell RNA-seq data into metabolic models to characterize tumour cell populations

MotivationMetabolic reprogramming is a general feature of cancer cells. Regrettably, the comprehensive quantification of metabolites in biological specimens does not promptly translate into knowledge on the utilization of metabolic pathways. Computational models hold the promise to bridge this gap, by estimating fluxes across metabolic pathways. Yet they currently portray the average behavior of intermixed subpopulations, masking their inherent heterogeneity known to hinder cancer diagnosis and treatment. If complemented with the information on single-cell transcriptome, now enabled by RNA sequencing (scRNA-seq), metabolic models of cancer populations are expected to empower the characterization of the mechanisms behind metabolic heterogeneity. To this aim, we propose single-cell Flux Balance Analysis (scFBA) as a computational framework to translate sc-transcriptomes into single-cell fluxomes.\n\nResultsWe show that the integration of scRNA-seq profiles of cells derived from lung ade-nocarcinoma and breast cancer patients, into a multi-scale stoichiometric model of cancer population: 1) significantly reduces the space of feasible single-cell fluxomes; 2) allows to identify clusters of cells with different growth rates within the population; 3) points out the possible metabolic interactions among cells via exchange of metabolites.\n\nAvailabilityThe scFBA suite of MATLAB functions is available at https://github.com/BIMIB-DISCo/scFBA, as well as the case study datasets.\n\nContactchiara.damiani@unimib.it

systems biology

Mechanisms of blood homeostasis: lineage tracking and a neutral model of cell populations in rhesus macaque

How a potentially diverse population of hematopoietic stem cells (HSCs) differentiates and proliferates to supply more than 1011 mature blood cells every day in humans remains a key biological question. We investigated this process by quantitatively analyzing the clonal structure of peripheral blood that is generated by a population of transplanted lentivirus-marked HSCs in myeloablated rhesus macaques. Each transplanted HSC generates a clonal lineage of cells in the peripheral blood that is then detected and quantified through deep sequencing of the viral vector integration sites (VIS) common within each lineage. This approach allowed us to observe, over a period of 4-12 years, hundreds of distinct clonal lineages. Surprisingly, while the distinct clone sizes varied by three orders of magnitude, we found that collectively, they form a steady-state clone size-distribution with a distinctive shape. Our concise model shows that slow HSC differentiation followed by fast progenitor growth is responsible for the observed broad clone size-distribution. Although all cells are assumed to be statistically identical, analogous to a neutral theory for the different clone lineages, our mathematical approach captures the intrinsic variability in the times to HSC differentiation after transplantation. Steady-state solutions of our model show that the predicted clone size-distribution is sensitive to only two combinations of parameters. By fitting the measured clone size-distributions to our mechanistic model, we estimate both the effective HSC differentiation rate and the number of active HSCs.

Systems Biology

Network Reconstruction from Perturbation Time Course Data

Networks underlie much of biology from subcellular to ecological scales. Yet, understanding what experimental data are needed and how to use them for unambiguously identifying the structure of even small networks remains a broad challenge. Here, we integrate a dynamic least squares framework into established modular response analysis (DL-MRA), that specifies sufficient experimental perturbation time course data to robustly infer arbitrary two and three node networks. DL-MRA considers important network properties that current methods often struggle to capture: (i) edge sign and directionality; (ii) cycles with feedback or feedforward loops including self-regulation; (iii) dynamic network behavior; (iv) edges external to the network; and (v) robust performance with experimental noise. We evaluate the performance of and the extent to which the approach applies to cell state transition networks, intracellular signaling networks, and gene regulatory networks. Although signaling networks are often an application of network reconstruction methods, the results suggest that only under quite restricted conditions can they be robustly inferred. For gene regulatory networks, the results suggest that incomplete knockdown is often more informative than full knockout perturbation, which may change experimental strategies for gene regulatory network reconstruction. Overall, the results give a rational basis to experimental data requirements for network reconstruction and can be applied to any such problem where perturbation time course experiments are possible.

systems biology

PIQED: Automated Identification And Quantification Of Protein Modifications From DIA-MS Data

Label-free quantification using data-independent acquisition (DIA) is a robust method for deep and accurate proteome quantification1,2. However, when lacking a pre-existing spectral library, as is often the case with studies of novel post-translational modifications (PTMs), samples are typically analyzed several times: one or more data dependent acquisitions (DDA) are used to generate a spectral library followed by DIA for quantification. This type of multi-injection analysis results in significant cost with regard to sample consumption and instrument time for each new PTM study, and may not be possible when sample amount is limiting and/or studies require a large number of biological replicates. Recently developed software (e.g. DIA-Umpire) has enabled combined peptide identification and quantification from a data-independent acquisition without any pre-existing spectral library3,4. Still, these tools are designed for protein level quantification. Here we demonstrate a software tool and workflow that extends DIA-Umpire to allow automated identification and quantification of PTM peptides from DIA. We accomplish this using a custom, open-source graphical user interface DIA-Pipe (https://github.com/jgmeyerucsd/PIQEDia/releases/tag/v0.1.2) (figure 1a).\n\nO_FIG O_LINKSMALLFIG WIDTH=195 HEIGHT=200 SRC=\"FIGDIR/small/141382_fig1.gif\" ALT=\"Figure 1\">\nView larger version (46K):\norg.highwire.dtl.DTLVardef@41b8f3org.highwire.dtl.DTLVardef@d59733org.highwire.dtl.DTLVardef@b9a722org.highwire.dtl.DTLVardef@8be50c_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1:C_FLOATNO Automated Qualitative and Quantitative Analysis of Post-Translational Modifications using DIA-Pipe. (a) Workflow of the all-DIA strategy for identification and quantification of PTMs. Modified peptides enriched from biological samples are analyzed by data-independent acquisition. All data analysis steps starting from instrument. wiff files, acquired on a TripleTOF 5600, can be completed using the DIA-Pipe GUI, including: (1) file conversion and pseudo-MS/MS spectra generation using DIA-Umpire, (2) database searching by MS-GF+, X! Tandem, and Comet followed by results refinement and combination using PeptideProphet/iProphet/PTMProphet, (3) automated spectral library generation and fragment area extraction using SkylineRunner, and finally, (4) Skyline report filtering and formatting for significance testing with mapDIA. (b) mProphet composite score distributions of target and second-best peaks picked by Skylineshowing essentially error-free peak picking by Skyline of peptides identified using pseudo-MS/MS spectra. (c) Distribution of coefficient of variations observed from three technical replicates for 1,182 acetylation sites. (d) Observed distributions of log2(fold change) computed by mapDIA using three technical replicates of 1X injection volume compared to three technical replicates of 0.5X injection volume; expected log2(fold change) value =1. (e) Number of identified peptides from each single replicate injection and from the combination of two or three replicates. An average of 1,655 acetylated peptides were identified per replicate. The combination of two or three replicates increased the number of identifications by 19% or 29%, respectively.\n\nC_FIG

systems biology

Identification of Gene Regulation Models from Single-Cell Data

In quantitative analyses of biological processes, one may use many different scales of models (e.g., spatial or non-spatial, deterministic or stochastic, time-varying or at steady-state) or many different approaches to match models to experimental data (e.g., model fitting or parameter uncertainty/sloppiness quantification with different experiment designs). These different analyses can lead to surprisingly different results, even when applied to the same data and the same model. We use a simplified gene regulation model to illustrate many of these concerns, especially for ODE analyses of deterministic processes, chemical master equation and finite state projection analyses of heterogeneous processes, and stochastic simulations. For each analysis, we employ MO_SCPLOWATLABC_SCPLOW and PO_SCPLOWYTHONC_SCPLOW software to consider a time-dependent input signal (e.g., a kinase nuclear translocation) and several model hypotheses, along with simulated single-cell data. We illustrate different approaches (e.g., deterministic and stochastic) to identify the mechanisms and parameters of the same model from the same simulated data. For each approach, we explore how uncertainty in parameter space varies with respect to the chosen analysis approach or specific experiment design. We conclude with a discussion of how our simulated results relate to the integration of experimental and computational investigations to explore signal-activated gene expression models in yeast [1] and human cells [2]{ddagger}.\n\nPACS numbers: 87.10.+e, 87.15.Aa, 05.10.Gg, 05.40.Ca,02.50.-r\n\nSubmitted to: Phys. Biol.

systems biology

Integration of Molecular Interactome and Targeted Interaction Analysis to Identify a COPD Disease Network Module

The polygenic nature of complex diseases offers potential opportunities to utilize network-based approaches that leverage the comprehensive set of protein-protein interactions (the human interactome) to identify new genes of interest and relevant biological pathways. However, the incompleteness of the current human interactome prevents it from reaching its full potential to extract network-based knowledge from gene discovery efforts, such as genome-wide association studies, for complex diseases like chronic obstructive pulmonary disease (COPD). Here, we provide a framework that integrates the existing human interactome information with new experimental protein-protein interaction data for FAM13A, one of the most highly associated genetic loci to COPD, to find a more comprehensive disease network module. We identified an initial disease network neighborhood by applying a random-walk method. Next, we developed a network-based closeness approach (CAB) that revealed 9 out of 96 FAM13A interacting partners identified by affinity purification assays were significantly close to the initial network neighborhood. Moreover, compared to a similar method (local radiality), the CAB approach predicts low-degree genes as potential candidates. The candidates identified by the network-based closeness approach were combined with the initial network neighborhood to build a comprehensive disease network module (163 genes) that was enriched with genes differentially expressed between controls and COPD subjects in alveolar macrophages, lung tissue, sputum, blood, and bronchial brushing datasets. Overall, we demonstrate an approach to find disease-related network components using new laboratory data to overcome incompleteness of the current interactome.

systems biology

Short CT-rich motifs can trigger context-specific silencing of gene expression in bacteria

We use an oligonucleotide library of over 10000 variants together with a synthetic biology approach to identify an insulation mechanism encoded within a subset of {sigma}54 promoters. Insulation manifests itself as dramatically reduced protein expression for a downstream gene that may be expressed by transcriptional read-through. The insulation we observe is strongly associated with the presence of short CT-rich motifs (3-5 bp), positioned within 25 bp upstream of the Shine-Dalgarno (SD) motif of the silenced gene. We hypothesize that insulation is effected by binding of the RBS to the upstream CT-rich motif. We provide evidence to support this hypothesis using mutations to the CT-rich motif and gene expression measurements on multiple sequence variants. Modelling is also consistent with this hypothesis. We show that the strength of the silencing, effected by insulation, depends on the location and number of CT-rich motifs encoded within the promoters. Finally, we show that in E.coli these insulator sequences are preferentially encoded within {sigma}54 promoters as compared to other promoter types, suggesting a regulatory role for these sequences in natural contexts. Our findings suggest that context-related regulatory effects may often be due to sequence-specific interactions encoded sparsely by short motifs that are not easily detected by lower throughput studies. Such short sequence-specific phenomena can be uncovered with a focused OL design that filters out the sequence noise, as exemplified herein.

systems biology

Optimal feedback mechanisms for regulating cell numbers

How living cells employ counting mechanisms to regulate their numbers or density is a long-standing problem in developmental biology that ties directly with organism or tissue size. Diverse cells types have been shown to regulate their numbers via secretion of factors in the extracellular space. These factors act as a proxy for the number of cells and function to reduce cellular proliferation rates creating a negative feedback. It is desirable that the production rate of such factors be kept as low as possible to minimize energy costs and detection by predators. Here we formulate a stochastic model of cell proliferation with feedback control via a secreted extracellular factor. Our results show that while low levels of feedback minimizes random fluctuations in cell numbers around a given set point, high levels of feedback amplify Poisson fluctuations in secreted-factor copy numbers. This trade-off results in an optimal feedback strength, and sets a fundamental limit to noise suppression in cell numbers. Intriguingly, this fundamental limit depends additively on two variables: relative half-life of the secreted factor with respect to the cell proliferation rate, and the average number of factors secreted in a cells lifespan. We further expand the model to consider external disturbances in key physiological parameters, such as, proliferation and factor synthesis rates. Intriguingly, while negative feedback effectively mitigates disturbances in the proliferation rate, it amplifies disturbances in the synthesis rate. In summary, these results provide unique insights into the functioning of feedback-based counting mechanisms, and apply to organisms ranging from unicellular prokaryotes and eukaryotes to human cells.

systems biology

Endogenous miRNA sponges mediate the generation of oscillatory dynamics for a non-coding RNA network

Oscillations are crucial to the sustenance of living organisms, across a wide variety of biological processes. In eukaryotes, oscillatory dynamics are thought to arise from interactions at the protein and RNA levels; however, the role of non-coding RNA in regulating these dynamics remains understudied. In this work, using a mathematical model, we show how non-coding RNA acting as microRNA (miRNA) sponges in a conserved miRNA - transcription factor feedback motif, can give rise to oscillatory behaviour. Control of these non-coding RNA can dynamically create oscillations or stability, and we show how this behaviour predisposes to oscillations in the stochastic limit. These results, supported by emerging evidence for the role of miRNA sponges in development, point towards key roles of different species of miRNA sponges, such as circular RNA, potentially in the maintenance of yet unexplained oscillatory behaviour. These results help to provide a paradigm for understanding functional differences between the many redundant, but distinct RNA species thought to act as miRNA sponges in nature, such as long non-coding RNA, pseudogenes, competing mRNA, circular RNA, and 3 UTRs.\n\nAuthor summaryWe analyze the effects of a newly discovered species of non-coding RNA, acting as microRNA (miRNA) sponges, on intracellular signalling dynamics. We show that oscillatory behaviour can arise in a time-varying manner in an over-represented transcriptional feedback network. These results point towards novel hypotheses for the roles of different species of miRNA sponges, such as their increasingly understood role in neural development.

systems biology

Non-latching positive feedback enables robust bimodality by de-coupling expression noise from the mean

Fundamental to biological decision-making is the ability to generate bimodal expression patterns where two alternate expression states simultaneously exist. Here, we use a combination of single-cell analysis and mathematical modeling to examine the sources of bimodality in the transcriptional program controlling HIVs fate decision between active replication and viral latency. We find that the HIV Tat protein manipulates the intrinsic toggling of HIVs promoter, the LTR, to generate bimodal ON-OFF expression, and that transcriptional positive feedback from Tat shifts and expands the regime of LTR bimodality. This result holds for both minimal synthetic viral circuits and full-length virus. Strikingly, computational analysis indicates that the Tat circuits non-cooperative non-latching feedback architecture is optimized to slow the promoters toggling and generate bimodality by stochastic extinction of Tat. In contrast to the standard Poisson model, theory and experiment show that non-latching positive feedback substantially dampens the inverse noise-mean relationship to maintain stochastic bimodality despite increasing mean-expression levels. Given the rapid evolution of HIV, the presence of a circuit optimized to robustly generate bimodal expression appears consistent with the hypothesis that HIVs decision between active replication and latency provides a viral fitness advantage. More broadly, the results suggest that positive-feedback circuits may have evolved not only for signal amplification but also for robustly generating bimodality by decoupling expression fluctuations (noise) from mean expression levels.

systems biology

Controllability in an islet specific regulatory network identifies the transcriptional factor NFATC4, which regulates Type 2 Diabetes associated genes

Probing the dynamic control features of biological networks represents a new frontier in capturing the dysregulated pathways in complex diseases. Here, using patient samples obtained from a pancreatic islet transplantation program, we constructed a tissue-specific gene regulatory network and used the control centrality (Cc) concept to identify the high control centrality (HiCc) pathways, which might serve as key pathobiological pathways for Type 2 Diabetes (T2D). We found that HiCc pathway genes were significantly enriched with modest GWAS p-values in the DIAbetes Genetics Replication And Meta-analysis (DIAGRAM) study. We identified variants regulating gene expression (expression quantitative loci, eQTL) of HiCc pathway genes in islet samples. These eQTL genes showed higher levels of differential expression compared to non-eQTL genes in low, medium and high glucose concentrations in rat islets. Among genes with highly significant eQTL evidence, NFATC4 belonged to four HiCc pathways. We asked if the expressions of T2D-associated candidate genes from GWAS and literature are regulated by Nfatc4 in rat islets. Extensive in vitro silencing of Nfatc4 in rat islet cells displayed reduced expression of 16, and increased expression of 4 putative downstream T2D genes. Overall, our approach uncovers the mechanistic connection of NFATC4 with downstream targets including a previously unknown one, TCF7L2, and establishes the HiCc pathways relationship to T2D.

systems biology

Numerical Analysis of the Immersed Boundary Method for Cell-Based Simulation

Mathematical modelling provides a useful framework within which to investigate the organization of biological tissues. With advances in experimental biology leading to increasingly detailed descriptions of cellular behaviour, models that consider cells as individual objects are becoming a common tool to study how processes at the single-cell level affect collective dynamics and determine tissue size, shape and function. However, there often remains no comprehensive account of these models, their method of solution, computational implementation or analysis of parameter scaling, hindering our ability to utilise and accurately compare different models. Here we present an effcient, open-source implementation of the immersed boundary method (IBM), tailored to simulate the dynamics of cell populations. This approach considers the dynamics of elastic membranes, representing cell boundaries, immersed in a viscous Newtonian fluid. The IBM enables complex and emergent cell shape dynamics, spatially heterogeneous cell properties and precise control of growth mechanisms. We solve the model numerically using an established algorithm, based on the fast Fourier transform, providing full details of all technical aspects of our implementation. The implementation is undertaken within Chaste, an open-source C++ library that allows one to easily change constitutive assumptions. Our implementation scales linearly with time step, and subquadratically with mesh spacing and immersed boundary node spacing. We identify the relationship between the immersed boundary node spacing and fluid mesh spacing required to ensure fluid volume conservation within immersed boundaries, and the scaling of cell membrane stiffness and cell-cell interaction strength required when refining the immersed boundary discretization. This study provides a recipe allowing consistent parametrisation of IBM models.

Systems Biology

Evolutionary tradeoffs and the structure of allelic polymorphisms

Populations of organisms show prevalent genetic differences called polymorphisms. Understanding the effects of polymorphisms is of central importance in biology and medicine. Here, we ask which polymorphisms occur at high frequency when organisms evolve under tradeoffs between multiple tasks. Multiple tasks present a problem, because it is not possible to be optimal at all tasks simultaneously and hence compromises are necessary. Recent work indicates that tradeoffs lead to a simple geometry of phenotypes in the space of traits: phenotypes fall on the Pareto front, which is shaped as a polytope: a line, triangle, tetrahedron etc. The vertices of these polytopes are the optimal phenotypes for a single task. Up to now, work on this Pareto approach has not considered its genetic underpinnings. Here, we address this by asking how the polymorphism structure of a population is affected by evolution under tradeoffs. We simulate a multi-task selection scenario, in which the population evolves to the Pareto front: the line segment between two archetypes or the triangle between three archetypes. We find that polymorphisms that become prevalent in the population have pleiotropic phenotypic effects that align with the Pareto front. Similarly, epistatic effects between prevalent polymorphisms are parallel to the front. Alignment with the front occurs also for asexual mating. Alignment is reduced when drift or linkage is strong, and is replaced by a more complex structure in which many perpendicular allele effects cancel out. Aligned polymorphism structure allows mating to produce offspring that stand a good chance of being optimal multi-taskers in at least one of the locales available to the species.

systems biology

A minimalistic resource allocation model to explain ubiquitous increase in protein expression with growth rate

Most proteins show changes in level across growth conditions. Many of these changes seem to be coordinated with the specific growth rate rather than the growth environment or the protein function. Although cellular growth rates, gene expression levels and gene regulation have been at the center of biological research for decades, there are only a few models giving a base line prediction of the dependence of the proteome fraction occupied by a gene with the specific growth rate.\n\nWe present a simple model that predicts a widely coordinated increase in the fraction of many proteins out of the proteome, proportionally with the growth rate. The model reveals how passive redistribution of resources, due to active regulation of only a few proteins, can have proteome wide effects that are quantitatively predictable. Our model provides a potential explanation for why and how such a coordinated response of a large fraction of the proteome to the specific growth rate arises under different environmental conditions. The simplicity of our model can also be useful by serving as a baseline null hypothesis in the search for active regulation. We exemplify the usage of the model by analyzing the relationship between growth rate and proteome composition for the model microorganism E.coli as reflected in two recent proteomics data sets spanning various growth conditions. We find that the fraction out of the proteome of a large number of proteins, and from different cellular processes, increases proportionally with the growth rate. Notably, ribosomal proteins, which have been previously reported to increase in fraction with growth rate, are only a small part of this group of proteins. We suggest that, although the fractions of many proteins change with the growth rate, such changes could be part of a global effect, not requiring specific cellular control mechanisms.

Systems Biology

Antisense transcription-dependent chromatin signature modulates sense transcription and transcript dynamics

Antisense transcription is widespread in genomes. Despite large differences in gene size and architecture, we find that yeast and human genes share a unique, antisense transcription-associated chromatin signature. We asked whether this signature is related to a biological function for antisense transcription. Using quantitative RNA-FISH, we observed changes in sense transcript distributions in nuclei and cytoplasm as antisense transcript levels were altered. To determine the mechanistic differences underlying these distributions, we developed a mathematical framework describing transcription from initiation to transcript degradation. At GAL1, high levels of antisense transcription alters sense transcription dynamics, reducing rates of transcript production and processing, while increasing transcript stability, which is also a genome-wide association. Establishing the antisense transcription-associated chromatin signature through disruption of the Set3C histone deacetylase activity is sufficient to similarly change these rates even in the absence of antisense transcription. Thus, antisense transcription alters sense transcription dynamics in a chromatin-dependent manner.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=198 HEIGHT=200 SRC=\"FIGDIR/small/187237_fig8.gif\" ALT=\"Figure 8\">\nView larger version (52K):\norg.highwire.dtl.DTLVardef@4f872corg.highwire.dtl.DTLVardef@1339662org.highwire.dtl.DTLVardef@1d62cc2org.highwire.dtl.DTLVardef@148f24_HPS_FORMAT_FIGEXP M_FIG Graphical Abstract\n\nC_FIG In this work, Brown et al. provide a mechanistic understanding of the effect of antisense transcription on the production and fate of sense transcripts. Antisense transcription buffers genes against the action of the Set3 lysine deacetylase, thus altering rates of transcript production, processing and stability. O_LIConserved antisense transcription-dependent chromatin architecture near promoters\nC_LIO_LIAntisense transcription alters sense transcription dynamics and transcript stability\nC_LIO_LIAntisense transcription functions in a chromatin-dependent manner\nC_LIO_LIIncreased acetylation by set3{Delta} mimics high antisense transcriptional dynamics\nC_LI

systems biology

Improving the consistency of functional genomics screens using molecular features - a multi-omics, pan-cancer study

Probing the genetic dependencies of cancer cells helps understand the tumor biology and identify potential drug targets. RNAi-based shRNA and CRISPR/Cas9-based sgRNA have been commonly utilized in functional genetic screens to identify essential genes affecting growth rates in cancer cell lines. However, questions remain whether the gene essentiality profiles determined using these two technologies are comparable. In the present study, we collected 42 cell lines representing a variety of 10 tissue types, which had been screened both by shRNA and CRISPR techniques. We observed poor consistency of the essentiality scores between the two screens for the majority of the cell lines. The consistency did not improve after correcting the off-target effects in the shRNA screening, suggesting a minimal impact of off-target effects. We considered a linear regression model where the shRNA essentiality score is the predictor and the CRISPR essentiality score is the response variable. We showed that by including molecular features such as mutation, gene expression and copy number variation as covariates, the predictability of the regression model greatly improved, suggesting that molecular features may provide critical information in explaining the discrepancy between the shRNA and CRISPR-based essentiality scores. We provided a Combined Essentiality Score (CES) based on the model prediction and showed that the CES greatly improved the consensus of common essential genes. Furthermore, the CES also identified novel essential genes that are specific to individual cell types. Taken together, we provided a systematic approach to define a more accurate gene essentiality profile by integrating functional screen data and molecular profiles.

systems biology