Search bioRxivSearch

Biology subjects

Quackenbush, J.

Publications and source records attributed to Quackenbush, J..

13 recordsLinked to original sources

Gene regulatory network analysis identifies sex-linked differences in colon cancer drug metabolism processes

Significant sex differences are observed in colon cancer, and understanding these differences is essential to advance disease prevention, diagnosis, and treatment. Males have a higher risk of developing colon cancer and a lower survival rate than women. However, the molecular features that drive these sex differences are poorly understood. We used both transcript-based and gene regulatory network methods to analyze RNA-Seq data from The Cancer Genome Atlas for 445 patients with colon cancer. We compared gene expression between tumors in men and women and found no significant sex differences except for sex-chromosome genes. We then inferred patient-specific gene regulatory networks, and found significant regulatory differences between males and females, with drug and xenobiotics metabolism via cytochrome P450 pathways more strongly targeted in females. This finding was validated in a dataset that included 1,193 patients from five independent studies. While targeting of the drug metabolism pathway did not change the overall survival for males treated with adjuvant chemotherapy, females with greater targeting had an increase in 10-year overall survival probability, with 89% (95% CI: 78%-100%) survival compared to 61% (95% CI: 45%-82%) for women with lower targeting, respectively (p=0.034). Our network analysis uncovered patterns of transcriptional regulation that differentiate male and female colon cancer. Most importantly, targeting of the drug metabolism pathway was predictive of survival in women who received adjuvant chemotherapy. This network-based approach can be used to investigate the molecular features that drive sex differences in other cancers and complex diseases.

systems biology

Cancer subtype identification using somatic mutation data

BACKGROUNDWith the onset of next generation sequencing technologies, we have made great progress in identifying recurrent mutational drivers of cancer. As cancer tissues are now frequently screened for specific sets of mutations, a large amount of samples has become available for analysis. Classification of patients with similar mutation profiles may help identifying subgroups of patients who might benefit from specific types of treatment. However, classification based on somatic mutations is challenging due to the sparseness and heterogeneity of the data.\n\nMETHODSHere, we describe a new method to de-sparsify somatic mutation data using biological pathways. We applied this method to 23 cancer types from The Cancer Genome Atlas, including samples from 5, 805 primary tumors.\n\nRESULTSWe show that, for most cancer types, de-sparsified mutation data associates with phenotypic data. We identify poor prognostic subtypes in three cancer types, which are associated with mutations in signal transduction pathways for which targeted treatment options are available. We identify subtype-drug associations for 14 additional subtypes. Finally, we perform a pan-cancer subtyping analysis and identify nine pan-cancer subtypes, which associate with mutations in four overarching sets of biological pathways.\n\nCONCLUSIONSThis study is an important step towards understanding mutational patterns in cancer.

genomics

WebMeV: A Cloud Platform for Analyzing and Visualizing Cancer Genomic Data

Although large, complex genomic data sets are increasingly easy to generate, and the number of publicly available data sets in cancer and other diseases is rapidly growing, the lack of intuitive, easy to use analysis tools has remained a barrier to the effective use of such data. WebMeV (https://mev.tm4.org) is an open-source, web-based tool that gives users access to sophisticated tools for analysis of RNA-Seq and other data in an interface designed to democratize data access. WebMeV combines cloud-based technologies with a simple user interface to allow users to access large public data sets such as that from The Cancer Genome Atlas (TCGA) or to upload their own. The interface allows users to visualize data and to apply advanced data mining analysis methods to explore the data and draw biologically meaningful conclusions. We provide an overview of WebMeV and demonstrate two simple use cases that illustrate the value of putting data analysis in the hands of those looking to explore the underlying biology of the systems being studied.

bioinformatics

Phenotype-Driven Transitions In Regulatory Network Structure

Complex traits and diseases like human height or cancer are often not caused by a single mutation or genetic variant, but instead arise from multiple factors that together functionally perturb the underlying molecular network. Biological networks are known to be highly modular and contain dense \"communities\" of genes that carry out cellular processes, but these structures change between tissues, during development, and in disease. While many methods exist for inferring networks, we lack robust methods for quantifying changes in network structure. Here, we describe ALPACA (ALtered Partitions Across Community Architectures), a method for comparing two genome-scale networks derived from different phenotypic states to identify condition-specific modules. In simulations, ALPACA leads to more nuanced, sensitive, and robust module discovery than currently available network comparison methods. We used ALPACA to compare transcriptional networks in three contexts: angiogenic and non-angiogenic subtypes of ovarian cancer, human fibroblasts expressing transforming viral oncogenes, and sexual dimorphism in human breast tissue. In each case, ALPACA identified modules enriched for processes relevant to the phenotype. For example, modules specific to angiogenic ovarian tumors were enriched for genes associated with blood vessel development, interferon signaling, and flavonoid biosynthesis. In comparing the modular structure of networks in female and male breast tissue, we found that female breast has distinct modules enriched for genes involved in estrogen receptor and ERK signaling. The functional relevance of these new modules indicate that not only does phenotypic change correlate with network structural changes, but also that ALPACA can identify such modules in complex networks.\n\nSignificance statementDistinct phenotypes are often thought of in terms of unique patterns of gene expression. But the expression levels of genes and proteins are driven by networks of interacting elements, and changes in expression are driven by changes in the structure of the associated networks. Because of the size and complexity of these networks, identifying functionally significant changes in network topology has been an ongoing challenge. We describe a new method for comparing networks derived from related conditions, such as healthy and disease tissue, and identifying emergent modules associated with the phenotypic differences between the conditions. We show that this method can find both known and previously unreported pathways involved in three contexts: ovarian cancer, tumor viruses, and breast tissue development.

systems biology

Challenges And Emerging Directions In Single-Cell Analysis

Single-cell analysis is a rapidly evolving approach to characterize genome-scale molecular information at the individual cell level. Development of single-cell technologies and computational methods has enabled systematic investigation of cellular heterogeneity in a wide range of tissues and cell populations, yielding fresh insights into the composition, dynamics, and regulatory mechanisms of cell states in development and disease. Despite substantial advances, significant challenges remain in the analysis, integration, and interpretation of single-cell omics data. Here, we discuss the state of the field and recent advances, and look to future opportunities.

genomics

Histopathological image QTL discovery of thyroid autoimmune disease variants

Genotype-to-phenotype association studies typically use macroscopic physiological measurements or molecular readouts as quantitative traits. There are comparatively few suitable quantitative traits available between cell and tissue length scales, a limitation that hinders our ability to identify variants affecting phenotype at many clinically informative levels. Here we show that quantitative image features, automatically extracted from histopathological imaging data, can be used for image Quantitative Trait Loci (iQTL) mapping and variant discovery. Using thyroid pathology images, clinical metadata, and genomics data from the Genotype and Tissue Expression (GTEx) project, we establish and validate a quantitative imaging biomarker for immune cell infiltration. A total of 100,215 variants were selected for iQTL profiling, and tested for genotype-phenotype associations with our quantitative imaging biomarker. Significant associations were found in HDAC9 and TXNDC5. We validated the TXNDC5 association using GTEx cis-expression QTL data, and an independent hypothyroidism dataset from the Electronic Medical Records and Genomics network.\n\nOne Sentence SummaryWe use a histopathological image QTL analysis to identify genomic variants associated with immune cell infiltration.

genomics

Understanding Tissue-specific Gene Regulation

Although all human tissues carry out common processes, tissues are distinguished by gene expres-sion patterns, implying that distinct regulatory programs control tissue-specificity. In this study, we investigate gene expression and regulation across 38 tissues profiled in the Genotype-Tissue Expression project. We find that network edges (transcription factor to target gene connections) have higher tissue-specificity than network nodes (genes) and that regulating nodes (transcription factors) are less likely to be expressed in a tissue-specific manner as compared to their targets (genes). Gene set enrichment analysis of network targeting also indicates that regulation of tissue-specific function is largely independent of transcription factor expression. In addition, tissue-specific genes are not highly targeted in their corresponding tissue-network. However, they do assume bottleneck positions due to variability in transcription factor targeting and the influence of non-canonical regulatory interactions. These results suggest that tissue-specificity is driven by context-dependent regulatory paths, providing transcriptional control of tissue-specific processes.

genomics

Estimating Drivers of Cell State Transitions Using Gene Regulatory Network Models

Specific cellular states are often associated with distinct gene expression patterns. These states are plastic, changing during development, or in the transition from health to disease. One relatively simple extension of this concept is to recognize that we can classify different cell-types by their active gene regulatory networks and that, consequently, transitions between cellular states can be modeled by changes in these underlying regulatory networks. Here we describe MONSTER, MOdeling Network State Transitions from Expression and Regulatory data, a regression-based method for inferring transcription factor drivers of cell state conditions at the gene regulatory network level. As a demonstration, we apply MONSTER to four different studies of chronic obstructive pulmonary disease to identify transcription factors that alter the network structure as the cell state progresses toward the disease-state. Our results demonstrate that MONSTER can find strong regulatory signals that persist across studies and tissues of the same disease and that are not detectable using conventional analysis methods based on differential expression. An R package implementing MONSTER is available at github.com/QuackenbushLab/MONSTER.

systems biology

A network-based approach to eQTL interpretation and SNP functional characterization

Expression quantitative trait locus (eQTL) analysis associates genotype with gene expression, but most eQTL studies only include cis-acting variants and generally examine a single tissue. We used data from 13 tissues obtained by the Genotype-Tissue Expression (GTEx) project v6.0 and, in each tissue, identified both cis- and trans-eQTLs. For each tissue, we represented significant associations between single nucleotide polymorphisms (SNPs) and genes as edges in a bipartite network. These networks are organized into dense, highly modular communities often representing coherent biological processes. Global network hubs are enriched in distal gene regulatory regions such as enhancers, but are devoid of disease-associated SNPs from genome wide association studies. In contrast, local, community-specific network hubs (core SNPs) are preferentially located in regulatory regions such as promoters and enhancers and highly enriched for trait and disease associations. These results provide help explain how many weak-effect SNPs might together influence cellular function and phenotype.

genomics

Smooth Quantile Normalization

Between-sample normalization is a critical step in genomic data analysis to remove systematic bias and unwanted technical variation in high-throughput data. Global normalization methods are based on the assumption that observed variability in global properties is due to technical reasons and are unrelated to the biology of interest. For example, some methods correct for differences in sequencing read counts by scaling features to have similar median values across samples, but these fail to reduce other forms of unwanted technical variation. Methods such as quantile normalization transform the statistical distributions across samples to be the same and assume global differences in the distribution are induced by only technical variation. However, it remains unclear how to proceed with normalization if these assumptions are violated, for example if there are global differences in the statistical distributions between biological conditions or groups, and external information, such as negative or control features, is not available. Here we introduce a generalization of quantile normalization, referred to as smooth quantile normalization (qsmooth), which is based on the assumption that the statistical distribution of each sample should be the same (or have the same distributional shape) within biological groups or conditions, but allowing that they may differ between groups. We illustrate the advantages of our method on several high-throughput datasets with global differences in distributions corresponding to different biological conditions. We also perform a Monte Carlo simulation study to illustrate the bias-variance tradeoff of qsmooth compared to other global normalization methods. A software implementation is available from https://github.com/stephaniehicks/qsmooth.

genomics

Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse data

Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. We find analysis of large RNA-Seq data sets requires both careful quality control and that one account for sparsity due to the heterogeneity intrinsic in multi-group studies. An R package instantiating our method for large-scale RNA-Seq normalization and preprocessing, YARN, is available at bioconductor.org/packages/yarn.\n\nHighlightsO_LIOverview of assumptions used in preprocessing and normalization\nC_LIO_LIPipeline for preprocessing, quality control, and normalization of large heterogeneous data\nC_LIO_LIA Bioconductor package for the YARN pipeline and easy manipulation of count data\nC_LIO_LIPreprocessed GTEx data set using the YARN pipeline available as a resource\nC_LI

bioinformatics

Transcriptional landscape of cell lines and their tissues of origin

Cell lines are an indispensable tool in biomedical research and often used as surrogates for tissues. An important question is how well a cell lines transcriptional and regulatory processes reflect those of its tissue of origin. We analyzed RNA-Seq data from GTEx for 127 paired Epstein-Barr virus transformed lymphoblastoid cell lines and whole blood samples; and 244 paired fibroblast cell lines and skin biopsies. A combination of gene expression and network analyses shows that while cell lines carry the expression signatures of their primary tissues, albeit at reduced levels, they also exhibit changes in their patterns of transcription factor regulation. Cell cycle genes are over-expressed in cell lines compared to primary tissue, and they have a reduction of repressive transcription factor targeting. Our results provide insight into the expression and regulatory alterations observed in cell lines and suggest that these changes should be considered when using cell lines as models.\n\nHighlightsO_LICell lines differ from their source tissues in gene expression and regulation\nC_LIO_LIDistinct cell lines share altered patterns of cell cycle regulation\nC_LIO_LICell cycle genes are less strongly targeted by repressive TFs in cell lines\nC_LIO_LICell lines share expression with their source tissue, but at reduced levels\nC_LI

genomics

Sexual dimorphism in gene expression and regulatory networks across human tissues

Sexual dimorphism manifests in many diseases and may drive sex-specific therapeutic responses. To understand the molecular basis of sexual dimorphism, we conducted a comprehensive assessment of gene expression and regulatory network modeling in 31 tissues using 8716 human transcriptomes from GTEx. We observed sexually dimorphic patterns of gene expression involving as many as 60% of autosomal genes, depending on the tissue. Interestingly, sex hormone receptors do not exhibit sexually dimorphic expression in most tissues; however, differential network targeting by hormone receptors and other transcription factors (TFs) captures their downstream sexually dimorphic gene expression. Furthermore, differential network wiring was found extensively in several tissues, particularly in brain, in which not all regions exhibit strong differential expression. This systems-based analysis provides a new perspective on the drivers of sexual dimorphism, one in which a repertoire of TFs plays important roles in sex-specific rewiring of gene regulatory networks.\n\nHighlightsO_LISexual dimorphism manifests in both gene expression and gene regulatory networks\nC_LIO_LISubstantial sexual dimorphism in regulatory networks was found in several tissues\nC_LIO_LIMany differentially regulated genes are not differentially expressed\nC_LIO_LISex hormone receptors do not exhibit sexually dimorphic expression in most tissues\nC_LI

genomics