Search bioRxivSearch

Biology subjects

Irizarry, R. A.

Publications and source records attributed to Irizarry, R. A..

11 recordsLinked to original sources

Interpretable Convolution Methods for Learning Genomic Sequence Motifs

The first-layer filters employed in convolutional neural networks tend to learn, or extract, spatial features from the data. Within their application to genomic sequence data, these learned features are often visualized and interpreted by converting them to sequence logos; an information-based representation of the consensus nucleotide motif. The process to obtain such motifs, however, is done through post-training procedures which often discard the filter weights themselves and instead rely upon finding those sequences maximally correlated with the given filter. Moreover, the filters collectively learn motifs with high redundancy, often simply shifted representations of the same sequence. We propose a schema to learn sequence motifs directly through weight constraints and transformations such that the individual weights comprising the filter are directly interpretable as either position weight matrices (PWMs) or information gain matrices (IGMs). We additionally leverage regularization to encourage learning highly-representative motifs with low inter-filter redundancy. Through learning PWMs and IGMs directly we present preliminary results showcasing how our method is capable of incorporating previously-annotated database motifs along with learning motifs de novo and then outline a pipeline for how these tools may be used jointly in a data application.

bioinformatics

Post-Hurricane Vital Statistics Expose Fragility of Puerto Rico’s Health System

ImportanceHurricane Maria made landfall in Puerto Rico on September 20, 2017. As recently as May of this year (2018), the official death count was 64. After a study describing a household survey reported a much higher death count estimate, as well as evidence of population displacement, extensive loss of services, and a prolonged death rate the government released death registry data. These newly released data will permit a better understanding of the effects of this hurricane.\n\nObjectiveProvide a detailed description of the effects on mortality of Hurricane Maria and compare to other hurricanes.\n\nDesignWe fit a statistical model to mortality data that accounts for seasonal and non-hurricane related yearly effects. We then estimated the deviation from the expected death rate as a function of time.\n\nSettingWe fit this model to 1985-2018 Puerto Rico daily data, which includes the dates of hurricanes Hugo, Georges, and Maria, 2015-2018 Florida daily data, which includes the dates of Hurricane Irma, 2002-2004 Louisiana monthly data, which includes the date of Hurricane Katrina, and 2000-2016 New Jersey monthly data, which includes the date of Hurricane Sandy.\n\nResultsWe find a prolonged increase in death rate after Maria and Katrina, lasting at least 207 and 125 days, resulting in excess deaths estimates of 3,400 (95% CI, 3,100-3,700), and 1,800 (95% CI, 1,600-2100) respectively, showing that Maria had a more long term damaging impact. Surprisingly, we also find that in 1998, Georges had a comparable impact to Katrinas with a prolonged increase of 106 days resulting in 1,400 (95% CI, 1,200-1,700) excess deaths. For Hurricane Maria, we find sharp increases in a small number of causes of deaths, including diseases of the circulatory, endocrine and respiratory system, as well as bacterial infections and suicides.\n\nConclusion and RelevanceOur analysis suggests that since at least 1998, Puerto Ricos health system has been in a precarious state. Without a substantial intervention, it appears that if hit with another strong hurricane, Puerto Ricans will suffer the unnecessary death of hundreds of its citizens.\n\nKey PointsQuestion: How does the effect of Hurricane Maria on mortality in Puerto Rico compare to the effect of other hurricanes in Puerto Rico and other United States jurisdictions?\n\nFindings: We estimate about 3,000 excess deaths after Maria, a higher toll than Katrina. Only other comparable effect was after Georges, also in Puerto Rico. For Georges and Maria, we observe a prolonged death rate increase of more than 10% lasting several months. The causes of death that increased after Maria are consistent with a collapsed health system\n\nMeaning: Puerto Ricos health system does not appear to be ready to withstand another strong hurricane.

epidemiology

Genome-wide repressive capacity of promoter DNA methylation is revealed through epigenomic manipulation

The scientific community is increasingly embracing open science. This growing commitment to open science should be applauded and encouraged, especially when it occurs voluntarily and prior to peer review. Thanks to other researchers dedication to open science, we have had the privilege of conducting a reanalysis of a landmark experiment published as a preprint with data made available in a public repository. The study in question found that promoter DNA methylation is frequently insufficient to induce transcriptional repression, which appears to contradict a large body of observational studies showing a strong association between DNA methylation and gene expression. This study was the first to evaluate whether forcibly methylating thousands of DNA promoter regions is sufficient to suppress gene expression. The authors data analysis did not find a strong relationship between promoter methylation and transcriptional repression. However, their analyses did not make full use of statistical inference and applied a normalization technique that removes global differences that are representative of the actual biological system. Here we reanalyze the data with an approach that includes statistical inference of differentially methylated regions, as well as a normalization technique that accounts for global expression differences. We find that forced DNA methylation of thousands of promoters overwhelmingly represses gene expression. In addition, we show that complementary epigenetic marks of active transcription are reduced as a result of DNA methylation. Finally, by studying whether these associations are sensitive to the CG density of promoters, we find no substantial differences in the association between promoters with and without a CG island. The code needed to reproduce are analysis is included in the public GitHub repository github.com/kdkorthauer/repressivecapacity.

genomics

Technology-independent estimation of cell type composition using differentially methylated regions

BackgroundHigh-resolution genome-wide measurement of DNA methylation (DNAm) has become a widely used assay in biomedical research. A major challenge in measuring DNAm is variability introduced from intra-sample cellular heterogeneity, which is a convolution of DNAm profiles across cell types. When this source of variability is confounded with an outcome of interest, if unaccounted for, false positives ensue. This is particularly problematic in epigenome-wide association studies for human disease performed on whole blood, a heterogeneous tissue. To account for this source of variability, a first step is to determine the actual cell proportions of each sample. Currently, the most effective approach is based on fitting a linear model in which one assumes the DNAm profiles of the representative cell types are known. However, we can only make this assumption when a dedicated experiment is performed to provide a plug-in estimate for these profiles. Although this method works well in practice, technology-specific biases lead to platform-dependent plug-in profiles. As a result, to apply the current methods across technologies we are required to repeat these costly experiments for each platform.\n\nResultsHere, we present a method that accurately estimates cell proportions agnostic to platform by first using experimental data to identify regions in which each cell type is clearly methylated or unmethlyated and model these as latent states. While the continuous measurements used in the linear model approaches are affected by platform-specific biases, the latent states are biologically driven and therefore technology independent, implying that experimental data only needs to be collected once. We demonstrate that our method accurately estimates the cell composition from whole blood samples and is applicable across multiple platforms, including microarray and sequencing platforms.

genomics

High-throughput identification of RNA nuclear enrichment sequences

One of the biggest surprises since the sequencing of the human genome has been the discovery of thousands of long noncoding RNAs (lncRNAs)1-6. Although lncRNAs and mRNAs are similar in many ways, they differ with lncRNAs being more nuclear-enriched and in several cases exclusively nuclear7,8. Yet, the RNA-based sequences that determine nuclear localization remain poorly understood9-11. Towards the goal of systematically dissecting the lncRNA sequences that impart nuclear localization, we developed a massively parallel reporter assay (MPRA). Unlike previous MPRAs12-15 that determine motifs important for transcriptional regulation, we have modified this approach to identify sequences sufficient for RNA nuclear enrichment for 38 human lncRNAs. Using this approach, we identified 109 unique, conserved nuclear enrichment regions, originating from 29 distinct lncRNAs. We also discovered two shorter motifs within our nuclear enrichment regions. We further validated the sufficiency of several regions to impart nuclear localization by single molecule RNA fluorescence in situ hybridization (smRNA-FISH). Taken together, these results provide a first systematic insight into the sequence elements responsible for the nuclear enrichment of lncRNA molecules.

genomics

Detection and accurate False Discovery Rate control of differentially methylated regions from Whole Genome Bisulfite Sequencing

With recent advances in sequencing technology, it is now feasible to measure DNA methylation at tens of millions of sites across the entire genome. In most applications, biologists are interested in detecting differentially methylated regions, composed of multiple sites with differing methylation levels among populations. However, current computational approaches for detecting such regions do not provide accurate statistical inference. A major challenge in reporting uncertainty is that a genome-wide scan is involved in detecting these regions, which needs to be accounted for. A further challenge is that sample sizes are limited due to the costs associated with the technology. We have developed a new approach that overcomes these challenges and assesses uncertainty for differentially methylated regions in a rigorous manner. Region-level statistics are obtained by fitting a generalized least squares (GLS) regression model with a nested autoregressive correlated error structure for the effect of interest on transformed methylation proportions. We develop an inferential approach, based on a pooled null distribution, that can be implemented even when as few as two samples per population are available. Here we demonstrate the advantages of our method using both experimental data and Monte Carlo simulation. We find that the new method improves the specificity and sensitivity of list of regions and accurately controls the False Discovery Rate (FDR).

genomics

Varying-Censoring Aware Matrix Factorization for Single Cell RNA-Sequencing

Single cell RNA-Seq (scRNA-Seq) has become the most widely used high-throughput technology for gene expression profiling of individual cells. The potential of being able to measure cell-to-cell variability at a high-dimensional genomic scale opens numerous new lines of investigation in basic and clinical research. For example, by identifying groups of cells with expression profiles unlike those observed in cells with known phenotypes, new cell types may be discovered. Dimension reduction followed by unsupervised clustering are the quantitative approaches typically used to facilitate such discoveries. However, a challenge for this approach is that most scRNA-Seq datasets are sparse, with the percentages of measurements reported as zero ranging from 35% to 99% across cells, and these zeros are partially explained by experimental inefficiencies that lead to censored data. Furthermore, the observed across-cell differences in the percentages of zeros are partly due to technical artifacts rather than biological differences. Unfortunately, standard dimension reduction approaches treat these censored values as true zeros, which leads to the identification of distorted low-dimensional factors. When these factors are used for clustering, the distortion leads to incorrect identification of biological groups. Here, we propose an approach that accounts for cell-specific censoring with a varying-censoring aware matrix factorization (VAMF) model that permits the identification of factors in the presence of the above described systematic bias. We demonstrate the advantages of our approach on published scRNA-Seq data and confirm these on simulated data.

genomics

Meta Analysis Of Microbiome Studies Identifies Shared And Disease-Specific Patterns

1Hundreds of clinical studies have been published that demonstrate associations between the human microbiome and a variety of diseases. Yet, fundamental questions remain on how we can generalize this knowledge. For example, if diseases are mainly characterized by a small number of pathogenic species, then new targeted antimicrobial therapies may be called for. Alternatively, if diseases are characterized by a lack of healthy commensal bacteria, then new probiotic therapies might be a better option. Results from individual studies, however, can be inconsistent or in conflict, and comparing published data is further complicated by the lack of standard processing and analysis methods.\n\nHere, we introduce the MicrobiomeHD database, which includes 29 published case-control gut microbiome studies spanning ten different diseases. Using standardized data processing and analyses, we perform a comprehensive crossdisease meta-analysis of these studies. We find consistent and specific patterns of disease-associated microbiome changes. A few diseases are associated with many individual bacterial associations, while most show only around 20 genus-level changes. Some diseases are marked by the presence of pathogenic microbes whereas others are characterized by a depletion of health-associated bacteria. Furthermore, over 60% of microbes associated with individual diseases fall into a set of \"core\" health and disease-associated microbes, which are associated with multiple disease states. This suggests a universal microbial response to disease.

microbiology

Accounting for GC-content bias reduces systematic errors and batch effects in ChIP-Seq peak callers

The main application of ChIP-seq technology is the detection of genomic regions that bind to a protein of interest. A large part of functional genomics public catalogs are based on ChIP-seq data. These catalogs rely on peak calling algorithms that infer protein-binding sites by detecting genomic regions associated with more mapped reads (coverage) than expected by chance as a result of the experimental protocol's lack of perfect specificity. We find that GC-content bias accounts for substantial variability in the observed coverage for ChIP-Seq experiments and that this variability leads to false-positive peak calls. More concerning is that GC-effect varies across experiments, with the effect strong enough to result in a substantial number of peaks called differently when different laboratories perform experiments on the same cell-line. However, accounting for GC-content in ChIP-Seq is challenging because the binding sites of interest tend to be more common in high GC-content regions, which confounds real biological signal with the unwanted variability. To account for this challenge we introduce a statistical approach that accounts for GC-effects on both non-specific noise and signal induced by the binding site. The method can be used to account for this bias in binding quantification as well to improve existing peak calling algorithms. We use this approach to show a reduction in false positive peaks as well as improved consistency across laboratories.

genomics

Smooth Quantile Normalization

Between-sample normalization is a critical step in genomic data analysis to remove systematic bias and unwanted technical variation in high-throughput data. Global normalization methods are based on the assumption that observed variability in global properties is due to technical reasons and are unrelated to the biology of interest. For example, some methods correct for differences in sequencing read counts by scaling features to have similar median values across samples, but these fail to reduce other forms of unwanted technical variation. Methods such as quantile normalization transform the statistical distributions across samples to be the same and assume global differences in the distribution are induced by only technical variation. However, it remains unclear how to proceed with normalization if these assumptions are violated, for example if there are global differences in the statistical distributions between biological conditions or groups, and external information, such as negative or control features, is not available. Here we introduce a generalization of quantile normalization, referred to as smooth quantile normalization (qsmooth), which is based on the assumption that the statistical distribution of each sample should be the same (or have the same distributional shape) within biological groups or conditions, but allowing that they may differ between groups. We illustrate the advantages of our method on several high-throughput datasets with global differences in distributions corresponding to different biological conditions. We also perform a Monte Carlo simulation study to illustrate the bias-variance tradeoff of qsmooth compared to other global normalization methods. A software implementation is available from https://github.com/stephaniehicks/qsmooth.

genomics

Selection Corrected Statistical Inference for Region Detection with High-throughput Assays

Scientists use high-dimensional measurement assays to detect and prioritize regions of strong signal in a spatially organized domain. Examples include finding methylation enriched genomic regions using microarrays and identifying active cortical areas using brain-imaging. The most common procedure for detecting potential regions is to group together neighboring sites where the signal passed a threshold. However, one needs to account for the selection bias induced by this opportunistic procedure to avoid diminishing effects when generalizing to a population. In this paper, we present a model and a method that permit population inference for these detected regions. In particular, we provide non-asymptotic point and confidence interval estimates for mean effect in the region, which account for the local selection bias and the non-stationary covariance that is typical of these data. Such summaries allow researchers to better compare regions of different sizes and different correlation structures. Inference is provided within a conditional one-parameter exponential family for each region, with truncations that match the constraints of selection. A secondary screening-and-adjustment step allows pruning the set of detected regions, while controlling the false-coverage rate for the set of regions that are reported. We illustrate the benefits of the method by applying it to detected genomic regions with differing DNA-methylation rates across tissue types. Our method is shown to provide superior power compared to non-parametric approaches.

bioinformatics