Search bioRxivSearch

Biology subjects

Paulson, J. N.

Publications and source records attributed to Paulson, J. N..

10 recordsLinked to original sources

An introduced crop plant is driving diversification of the virulent bacterial pathogen Erwinia tracheiphila

Erwinia tracheiphila is the causal agent of bacterial wilt of cucurbits, an economically important phytopathogen affecting few cultivated Cucurbitaceae host plant species in temperate Eastern North America. However, essentially nothing is known about E. tracheiphila population structure or genetic diversity. To address this shortcoming, a representative collection of 88 E. tracheiphila isolates was gathered from throughout its geographic range, and their genomes were sequenced. Phylogenomic analysis revealed three genetic clusters with distinct hrpT3SS virulence gene repertoires, host plant association patterns, and geographic distributions. The low genetic variation within each cluster suggests a recent population bottleneck followed by population expansion. We showed that in the field and greenhouse, cucumber (Cucumis sativus), which was introduced to North America by early Spanish conquistadors, is the most susceptible host plant species, and the only species susceptible to isolates from all three lineages. The establishment of large agricultural populations of highly susceptible C. sativus in temperate Eastern North America may have facilitated the original emergence of E. tracheiphila into cucurbit agro-ecosystems, and this introduced plant species may now be acting as a highly susceptible reservoir host. Our findings have broad implications for agricultural sustainability by drawing attention to how worldwide crop plant movement, agricultural intensification and locally unique environments may affect the emergence, evolution, and epidemic persistence of virulent microbial pathogens.\n\nImportanceErwinia tracheiphila is a virulent phytopathogen that infects two genera of cucurbit crop plants, Cucurbita spp. (pumpkin and squash) and Cucumis spp. (muskmelon and cucumber). One of the unusual ecological traits of this pathogen is that it is limited to temperate Eastern North America. Here, we complete the first large-scale sequencing of an E. tracheiphila isolate collection. From phylogenomic, comparative genomic, and empirical analyses, we find that introduced Cucumis spp. crop plants are driving the diversification of E. tracheiphila into multiple, closely related lineages. Together, the results from this study show that locally unique biotic (plant population) and abiotic (climate) conditions can drive the evolutionary trajectories of locally endemic pathogens in unexpected ways.

evolutionary biology

Cancer subtype identification using somatic mutation data

BACKGROUNDWith the onset of next generation sequencing technologies, we have made great progress in identifying recurrent mutational drivers of cancer. As cancer tissues are now frequently screened for specific sets of mutations, a large amount of samples has become available for analysis. Classification of patients with similar mutation profiles may help identifying subgroups of patients who might benefit from specific types of treatment. However, classification based on somatic mutations is challenging due to the sparseness and heterogeneity of the data.\n\nMETHODSHere, we describe a new method to de-sparsify somatic mutation data using biological pathways. We applied this method to 23 cancer types from The Cancer Genome Atlas, including samples from 5, 805 primary tumors.\n\nRESULTSWe show that, for most cancer types, de-sparsified mutation data associates with phenotypic data. We identify poor prognostic subtypes in three cancer types, which are associated with mutations in signal transduction pathways for which targeted treatment options are available. We identify subtype-drug associations for 14 additional subtypes. Finally, we perform a pan-cancer subtyping analysis and identify nine pan-cancer subtypes, which associate with mutations in four overarching sets of biological pathways.\n\nCONCLUSIONSThis study is an important step towards understanding mutational patterns in cancer.

genomics

Histopathological image QTL discovery of thyroid autoimmune disease variants

Genotype-to-phenotype association studies typically use macroscopic physiological measurements or molecular readouts as quantitative traits. There are comparatively few suitable quantitative traits available between cell and tissue length scales, a limitation that hinders our ability to identify variants affecting phenotype at many clinically informative levels. Here we show that quantitative image features, automatically extracted from histopathological imaging data, can be used for image Quantitative Trait Loci (iQTL) mapping and variant discovery. Using thyroid pathology images, clinical metadata, and genomics data from the Genotype and Tissue Expression (GTEx) project, we establish and validate a quantitative imaging biomarker for immune cell infiltration. A total of 100,215 variants were selected for iQTL profiling, and tested for genotype-phenotype associations with our quantitative imaging biomarker. Significant associations were found in HDAC9 and TXNDC5. We validated the TXNDC5 association using GTEx cis-expression QTL data, and an independent hypothyroidism dataset from the Electronic Medical Records and Genomics network.\n\nOne Sentence SummaryWe use a histopathological image QTL analysis to identify genomic variants associated with immune cell infiltration.

genomics

Understanding Tissue-specific Gene Regulation

Although all human tissues carry out common processes, tissues are distinguished by gene expres-sion patterns, implying that distinct regulatory programs control tissue-specificity. In this study, we investigate gene expression and regulation across 38 tissues profiled in the Genotype-Tissue Expression project. We find that network edges (transcription factor to target gene connections) have higher tissue-specificity than network nodes (genes) and that regulating nodes (transcription factors) are less likely to be expressed in a tissue-specific manner as compared to their targets (genes). Gene set enrichment analysis of network targeting also indicates that regulation of tissue-specific function is largely independent of transcription factor expression. In addition, tissue-specific genes are not highly targeted in their corresponding tissue-network. However, they do assume bottleneck positions due to variability in transcription factor targeting and the influence of non-canonical regulatory interactions. These results suggest that tissue-specificity is driven by context-dependent regulatory paths, providing transcriptional control of tissue-specific processes.

genomics

Metaviz: interactive statistical and visual analysis of metagenomic data

Along with the survey techniques of 16S rRNA amplicon and whole-metagenome shotgun sequencing, an array of tools exists for clustering, taxonomic annotation, normalization, and statistical analysis of microbiome sequencing results. Integrative and interactive visualization that enables researchers to perform exploratory analysis in this feature rich hierarchical data is an area of need. In this work, we present Metaviz, a web browser-based tool for interactive exploratory metagenomic data analysis. Metaviz can visualize abundance data served from an R session or a Python web service that queries a graph database. As metagenomic sequencing features have a hierarchy, we designed a novel navigation mechanism to explore this feature space. We visualize abundance counts with heatmaps and stacked bar plots that are dynamically updated as a user selects taxonomic features to inspect. Metaviz also supports common data exploration techniques, including PCA scatter plots to interpret variability in the dataset and alpha diversity boxplots for examining ecological community composition. The Metaviz application and documentation is hosted at http://www.metaviz.org.

bioinformatics

Longitudinal differential abundance analysis of microbial marker-gene surveys using smoothing splines

BackgroundHigh-throughput targeted sequencing of the 16S ribosomal RNA marker gene is often used to profile and characterize the taxonomic composition of microbial communities. This type of big high-through sequencing data is rapidly being applied to various infectious diseases like diarrhea. While many studies are limited to single \"snapshots\" of these communities, there is increasing recognition that longitudinal profiling of these communities are required to understand community dynamics and the complex relationships between dynamics and phenotypes of interest. Statistical methods that determine microbial features that are differentially expressed are required as an initial step to characterizing phenotypic associations with community dynamics in big data and infectious diseases.\n\nResultsWe present a novel method for longitudinal marker-gene surveys based on smoothing splines that allows discovery and inference of time periods where specific microbial features are differentially abundant. We applied our method to three 16S marker-gene surveys, including, groups of gnotobiotic mice on two diets, patients challenged with ETEC (H10407), and a vaginal microbiome of healthy women. Employing our methodology we recover known bacterial differences and highlight a few extra species providing insight into when specific changes occurred. Additionally, in the cohort challenged with ETEC we recover proposed probiotic bacteria Bacteroides xylanisolvens, Collinsella aerofaciens, and Faecalibacterium prausnitzii associatons with healthy individuals.\n\nConclusionsThe method presented is, to our knowledge, the first flexible method of its kind implemented as a software capable of detecting time periods of differential abundance for microbial features species between two or more sample groups of interest. Our method is available within the metagenomeSeq open-source software for analysis of metagenomic package available through the Bioconductor project and is termed metaSplines.

genomics

A network-based approach to eQTL interpretation and SNP functional characterization

Expression quantitative trait locus (eQTL) analysis associates genotype with gene expression, but most eQTL studies only include cis-acting variants and generally examine a single tissue. We used data from 13 tissues obtained by the Genotype-Tissue Expression (GTEx) project v6.0 and, in each tissue, identified both cis- and trans-eQTLs. For each tissue, we represented significant associations between single nucleotide polymorphisms (SNPs) and genes as edges in a bipartite network. These networks are organized into dense, highly modular communities often representing coherent biological processes. Global network hubs are enriched in distal gene regulatory regions such as enhancers, but are devoid of disease-associated SNPs from genome wide association studies. In contrast, local, community-specific network hubs (core SNPs) are preferentially located in regulatory regions such as promoters and enhancers and highly enriched for trait and disease associations. These results provide help explain how many weak-effect SNPs might together influence cellular function and phenotype.

genomics

Smooth Quantile Normalization

Between-sample normalization is a critical step in genomic data analysis to remove systematic bias and unwanted technical variation in high-throughput data. Global normalization methods are based on the assumption that observed variability in global properties is due to technical reasons and are unrelated to the biology of interest. For example, some methods correct for differences in sequencing read counts by scaling features to have similar median values across samples, but these fail to reduce other forms of unwanted technical variation. Methods such as quantile normalization transform the statistical distributions across samples to be the same and assume global differences in the distribution are induced by only technical variation. However, it remains unclear how to proceed with normalization if these assumptions are violated, for example if there are global differences in the statistical distributions between biological conditions or groups, and external information, such as negative or control features, is not available. Here we introduce a generalization of quantile normalization, referred to as smooth quantile normalization (qsmooth), which is based on the assumption that the statistical distribution of each sample should be the same (or have the same distributional shape) within biological groups or conditions, but allowing that they may differ between groups. We illustrate the advantages of our method on several high-throughput datasets with global differences in distributions corresponding to different biological conditions. We also perform a Monte Carlo simulation study to illustrate the bias-variance tradeoff of qsmooth compared to other global normalization methods. A software implementation is available from https://github.com/stephaniehicks/qsmooth.

genomics

Transcriptional landscape of cell lines and their tissues of origin

Cell lines are an indispensable tool in biomedical research and often used as surrogates for tissues. An important question is how well a cell lines transcriptional and regulatory processes reflect those of its tissue of origin. We analyzed RNA-Seq data from GTEx for 127 paired Epstein-Barr virus transformed lymphoblastoid cell lines and whole blood samples; and 244 paired fibroblast cell lines and skin biopsies. A combination of gene expression and network analyses shows that while cell lines carry the expression signatures of their primary tissues, albeit at reduced levels, they also exhibit changes in their patterns of transcription factor regulation. Cell cycle genes are over-expressed in cell lines compared to primary tissue, and they have a reduction of repressive transcription factor targeting. Our results provide insight into the expression and regulatory alterations observed in cell lines and suggest that these changes should be considered when using cell lines as models.\n\nHighlightsO_LICell lines differ from their source tissues in gene expression and regulation\nC_LIO_LIDistinct cell lines share altered patterns of cell cycle regulation\nC_LIO_LICell cycle genes are less strongly targeted by repressive TFs in cell lines\nC_LIO_LICell lines share expression with their source tissue, but at reduced levels\nC_LI

genomics

Sexual dimorphism in gene expression and regulatory networks across human tissues

Sexual dimorphism manifests in many diseases and may drive sex-specific therapeutic responses. To understand the molecular basis of sexual dimorphism, we conducted a comprehensive assessment of gene expression and regulatory network modeling in 31 tissues using 8716 human transcriptomes from GTEx. We observed sexually dimorphic patterns of gene expression involving as many as 60% of autosomal genes, depending on the tissue. Interestingly, sex hormone receptors do not exhibit sexually dimorphic expression in most tissues; however, differential network targeting by hormone receptors and other transcription factors (TFs) captures their downstream sexually dimorphic gene expression. Furthermore, differential network wiring was found extensively in several tissues, particularly in brain, in which not all regions exhibit strong differential expression. This systems-based analysis provides a new perspective on the drivers of sexual dimorphism, one in which a repertoire of TFs plays important roles in sex-specific rewiring of gene regulatory networks.\n\nHighlightsO_LISexual dimorphism manifests in both gene expression and gene regulatory networks\nC_LIO_LISubstantial sexual dimorphism in regulatory networks was found in several tissues\nC_LIO_LIMany differentially regulated genes are not differentially expressed\nC_LIO_LISex hormone receptors do not exhibit sexually dimorphic expression in most tissues\nC_LI

genomics