Search bioRxivSearch

Biology subjects

Robinson, M. D.

Publications and source records attributed to Robinson, M. D..

9 recordsLinked to original sources

A junction coverage compatibility score to quantify the reliability of transcript abundance estimates and annotation catalogs

Most methods for statistical analysis of RNA-seq data take a matrix of abundance estimates for some type of genomic features as their input, and consequently the quality of any obtained results are directly dependent on the quality of these abundances. Here, we present the junction coverage compatibility (JCC) score, which provides a way to evaluate the reliability of transcript-level abundance estimates as well as the accuracy of transcript annotation catalogs. It works by comparing the observed number of reads spanning each annotated splice junction in a genomic region to the predicted number of junction-spanning reads, inferred from the estimated transcript abundances and the genomic coordinates of the corresponding annotated transcripts. We show that while most genes show good agreement between the observed and predicted junction coverages, there is a small set of genes that do not. Genes with poor agreement are found regardless of the method used to estimate transcript abundances, and the corresponding transcript abundances should be treated with care in any downstream analyses.

bioinformatics

The proto CpG island methylator phenotype of sessile serrated adenoma/polyps

Sessile serrated adenomas/polyps (SSA/Ps) are the putative precursors of the {small tilde}20% of colon cancers with the CpG island methylator phenotype (CIMP), but their molecular features are poorly understood. We used high-throughput analysis of DNA methylation and gene expression to investigate the epigenetic phenotype of SSA/Ps. Fresh-tissue samples of 17 SSA/Ps and (for comparison purposes) 15 conventional adenomas (cADNs)--each with a matched sample of normal mucosa-- were prospectively collected during colonoscopy (total no. samples analyzed: 64). DNA and RNA were extracted from each sample. DNA was subjected to bisulfite next-generation sequencing to assess methylation levels at {small tilde}2.7 million CpG sites located predominantly in gene regulatory regions and spanning 80.5Mb ({small tilde}2.5% of the genome); RNA was sequenced to define the samples transcriptomes. An independent series of 61 archival lesions was used for targeted verification of DNA methylation findings. Compared with normal mucosa samples, SSA/Ps and cADNs exhibited markedly remodeled methylomes. In cADNs, hypomethylated regions were far more numerous (18,417 vs 4288 in SSA/Ps) and rarely affected CpG islands/shores. SSA/Ps seemed to have escaped this wave of demethylation. Cytosine hypermethylation in SSA/Ps was more pervasive (hypermethylated regions: 22,147 vs 15,965 in cADNs; hypermethylated genes: 4938 vs 3443 in cADNs) and more extensive (region for region), and it occurred mainly within CpG islands and shores. Given its resemblance to the CIMP typical of SSA/Ps putative descendant colon cancers, we refer to the SSA/P methylation phenotype as proto-CIMP. Verification studies of six hypermethylated regions (3 SSA/P-specific and 3 common) demonstrated the high potential of DNA methylation markers for predicting the diagnosis of SSA/Ps and cADNs. Surprisingly, proto-CIMP in SSA/Ps was associated with upregulated gene expression (n=618 genes vs 349 that were downregulated); downregulation was more common in cADNs (n=712 vs 516 upregulated genes). The epigenetic landscape of SSA/Ps differs markedly from that of cADNs. These differences are a potentially rich source of novel tissue-based and noninvasive biomarkers that can add precision to the clinical management of the two most frequent colon-cancer precursors.

cancer biology

diffcyt: Differential discovery in high-dimensional cytometry via high-resolution clustering

1High-dimensional flow and mass cytometry allow cell types and states to be characterized in great detail by measuring expression levels of more than 40 targeted protein markers per cell. Here, we present diffcyt, a new computational framework for differential discovery analyses in these datasets, based on (i) high-resolution clustering and (ii) empirical Bayes moderated tests adapted from transcriptomics. Our approach provides improved statistical performance, including for rare cell populations, along with flexible experimental designs and fast runtimes in an open-source framework.

bioinformatics

Observation weights to unlock bulk RNA-seq tools for zero inflation and single-cell applications

Dropout events in single-cell transcriptome sequencing (scRNA-seq) cause many transcripts to go undetected and induce an excess of zero read counts, leading to power issues in differential expression (DE) analysis. This has triggered the development of bespoke scRNA-seq DE methods to cope with zero inflation. Recent evaluations, however, have shown that dedicated scRNA-seq tools provide no advantage compared to traditional bulk RNA-seq tools. We introduce a weighting strategy, based on a zero-inflated negative binomial (ZINB) model, that identifies excess zero counts and generates gene and cell-specific weights to unlock bulk RNA-seq DE pipelines for zero-inflated data, boosting performance for scRNA-seq.

bioinformatics

Channel crosstalk correction in suspension and imaging mass cytometry

Mass cytometry enables simultaneous analysis of over 40 proteins and their modifications in single cells through use of metal-tagged antibodies. Compared to fluorescent dyes, the use of pure metal isotopes strongly reduces spectral overlap among measurement channels. Crosstalk still exists, however, caused by isotopic impurity, oxide formation, and mass cytometer properties. Spillover effects can be minimized, but not avoided, by following a set of constraining rules when designing an antibody panel. Generation of such low crosstalk panels requires considerable expert knowledge, knowledge of the abundance of each marker and substantial experimental effort. Here we describe a novel bead-based compensation workflow that includes R-based software and a web tool, which enables correction for interference between channels. We demonstrate utility in suspension mass cytometry and show how this approach can be applied to imaging mass cytometry. Our approach greatly simplifies the development of new antibody panels, increases flexibility for antibody-metal pairing, improves overall data quality, thereby reducing the risk of reporting cell phenotype and function artifacts, and greatly facilitates analysis of complex samples for which antigen abundances are unknown.

bioinformatics

zingeR: unlocking RNA-seq tools for zero-inflation and single cell applications

Dropout in single cell RNA-seq (scRNA-seq) applications causes many transcripts to go undetected. It induces excess zero counts, which leads to power issues in differential expression (DE) analysis and has triggered the development of bespoke scRNA-seq DE tools that cope with zero-inflation. Recent evaluations, however, have shown that dedicated scRNA-seq tools provide no advantage compared to traditional bulk RNA-seq tools. We introduce zingeR, a zero-inflated negative binomial model that identifies excess zero counts and generates observation weights to unlock bulk RNA-seq pipelines for zero-inflation, boosting performance in scRNA-seq differential expression analysis.

bioinformatics

An integrative strategy to identify the entire protein coding potential of prokaryotic genomes by proteogenomics

Accurate annotation of all protein-coding sequences (CDSs) is an essential prerequisite to fully exploit the rapidly growing repertoire of completely sequenced prokaryotic genomes. However, large discrepancies among the number of CDSs annotated by different resources, missed functional short open reading frames (sORFs), and overprediction of spurious ORFs represent serious limitations.\n\nOur strategy towards accurate and complete genome annotation consolidates CDSs from multiple reference annotation resources, ab initio gene prediction algorithms and in silico ORFs in an integrated proteogenomics database (iPtgxDB) that covers the entire protein-coding potential of a prokaryotic genome. By extending the PeptideClassifier concept of unambiguous peptides for prokaryotes, close to 95% of the identifiable peptides imply one distinct protein, largely simplifying downstream analysis. Searching a comprehensive Bartonella henselae proteomics dataset against such an iPtgxDB allowed us to unambiguously identify novel ORFs uniquely predicted by each resource, including lipoproteins, differentially expressed and membrane-localized proteins, novel start sites and wrongly annotated pseudogenes. Most novelties were confirmed by targeted, parallel reaction monitoring mass spectrometry, including unique ORFs and variants identified in a re-sequenced laboratory strain that are not present in its reference genome. We demonstrate the general applicability of our strategy for genomes with varying GC content and distinct taxonomic origin, and release iPtgxDBs for B. henselae, Bradyrhozibium diazoefficiens and Escherichia coli as well as the software to generate such proteogenomics search databases for any prokaryote.

genomics

Bias, Robustness And Scalability In Differential Expression Analysis Of Single-Cell RNA-Seq Data

BackgroundAs single-cell RNA-seq (scRNA-seq) is becoming increasingly common, the amount of publicly available data grows rapidly, generating a useful resource for computational method development and extension of published results. Although processed data matrices are typically made available in public repositories, the procedure to obtain these varies widely between data sets, which may complicate reuse and cross-data set comparison. Moreover, while many statistical methods for performing differential expression analysis of scRNA-seq data are becoming available, their relative merits and the performance compared to methods developed for bulk RNA-seq data are not sufficiently well understood.\n\nResultsWe present conquer, a collection of consistently processed, analysis-ready public single-cell RNA-seq data sets. Each data set has count and transcripts per million (TPM) estimates for genes and transcripts, as well as quality control and exploratory analysis reports. We use a subset of the data sets available in conquer to perform an extensive evaluation of the performance and characteristics of statistical methods for differential gene expression analysis, evaluating a total of 30 statistical approaches on both experimental and simulated scRNA-seq data.\n\nConclusionsConsiderable differences are found between the methods in terms of the number and characteristics of the genes that are called differentially expressed. Pre-filtering of lowly expressed genes can have important effects on the results, particularly for some of the methods originally developed for analysis of bulk RNA-seq data. Generally, however, methods developed for bulk RNA-seq analysis do not perform notably worse than those developed specifically for scRNA-seq.

bioinformatics

A general and powerful stage-wise testing procedure for differential expression and differential transcript usage

BackgroundReductions in sequencing cost and innovations in expression quantification have prompted an emergence of RNA-seq studies with complex designs and data analysis at transcript resolution. These applications involve multiple hypotheses per gene, leading to challenging multiple testing problems. Conventional approaches provide separate top-lists for every contrast and false discovery rate (FDR) control at individual hypothesis level. Hence, they fail to establish proper gene-level error control, which compromises downstream validation experiments. Tests that aggregate individual hypotheses are more powerful and provide gene-level FDR control, but in the RNA-seq literature no methods are available for post-hoc analysis of individual hypotheses.\n\nResultsWe introduce a two-stage procedure that leverages the increased power of aggregated hypothesis tests while maintaining high biological resolution by post-hoc analysis of genes passing the screening hypothesis. Our method is evaluated on simulated and real RNA-seq experiments. It provides gene-level FDR control in studies with complex designs while boosting power for interaction effects without compromising the discovery of main effects. In a differential transcript usage/expression context, stage-wise testing gains power by aggregating hypotheses at the gene level, while providing transcript-level assessment of genes passing the screening stage. Finally, a prostate cancer case study highlights the relevance of combining gene with transcript level results.\n\nConclusionStage-wise testing is a general paradigm that can be adopted whenever individual hypotheses can be aggregated. In our context, it achieves an optimal middle ground between biological resolution and statistical power while providing gene-level FDR control, which is beneficial for downstream biological interpretation and validation.

bioinformatics