Search bioRxivSearch

Biology subjects

Weller, J.

Publications and source records attributed to Weller, J..

3 recordsLinked to original sources

Transcriptomics data availability and reusability in the transition from microarray to next-generation sequencing

Over the last two decades, molecular biology has been changed by the introduction of high-throughput technologies. Data sharing requirements have prompted the establishment of persistent data archives. A standardized approach for recording and managing these data was first proposed in the Minimal Information About a Microarray Experiment (MIAME) guidelines. The Minimal Information about a high throughput nucleotide Sequencing Experiment (MINSEQE) proposal was introduced in 2008 as a logical extension of the guidelines to next-generation sequencing (NGS) technologies used for transcriptome analysis. We present a historical snapshot of the data-sharing situation focusing on transcriptomics data from both microarray and RNA-sequencing experiments published between 2009 and 2013, a period during which RNA-seq studies became increasingly popular for transcriptome analysis. We assess how much data from RNA-seq based experiments is actually available in persistent data archives, compared to data derived from microarray based experiments, and evaluate how these types of data differ. Based on this analysis, we provide recommendations to improve RNA-seq data availability, reusability, and reproducibility.

bioinformatics

Efficient design of maximally active and specific nucleic acid diagnostics for thousands of viruses

Diagnostics, particularly for rapidly evolving viruses, stand to benefit from a principled, measurement-driven design that harnesses machine learning and vast genomic data--yet the capability for such design has not been previously built. Here, we develop and extensively validate an approach to designing viral diagnostics that applies a learned model within a combinatorial optimization framework. Concentrating on CRISPR-based diagnostics, we screen a library of 19,209 diagnostic-target pairs and train a deep neural network that predicts, from RNA sequence alone, diagnostic signal better than contemporary techniques. Our model then makes it possible to design assays that are maximally sensitive over the spectrum of a viruss genomic variation. We introduce ADAPT (https://adapt.guide), a system for fully-automated design, and use ADAPT to design optimal diagnostics for the 1,933 vertebrate-infecting viral species within 2 hours for most species and 24 hours for all but 3. We experimentally show ADAPTs designs are sensitive and specific down to the lineage level, including against viruses that pose challenges involving genomic variation and specificity. ADAPTs designs exhibit significantly higher fluorescence and permit lower limits of detection, across a viruss entire variation, than the outputs of standard design techniques. Our model-based optimization strategy has applications broadly to viral nucleic acid diagnostics and other sequence-based technologies, and, paired with clinical validation, could enable a critically-needed, proactive resource of assays for surveilling and responding to pathogens.

genomics

AutoRELACS: Automated Generation And Analysis Of Ultra-parallel ChIP-seq

Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is a method used to profile protein-DNA interactions genome-wide. RELACS (Restriction Enzyme-based Labeling of Chromatin in Situ) is a recently developed ChIP-seq protocol that deploys a chromatin barcoding strategy to enable standardized and high-throughput generation of ChIP-seq data. The manual implementation of RELACS is constrained by human processivity in both data generation and data analysis. To overcome these limitations, we have developed AutoRELACS, an automated implementation of the RELACS protocol using the liquid handler Biomek i7 workstation. We match the unprecedented processivity in data generation allowed by AutoRELACS with the automated computation pipelines offered by snakePipes. In doing so, we build a continuous workflow that streamlines epigenetic profiling, from sample collection to biological interpretation. Here, we show that AutoRELACS successfully automates chromatin barcode integration, and is able to generate high-quality ChIP-seq data comparable with the standards of the manual protocol, also for limited amounts of biological samples.

genomics