Search bioRxivSearch

Biology subjects

Le Cao, K.-A.

Publications and source records attributed to Le Cao, K.-A..

3 recordsLinked to original sources

scRNA-seq mixology: towards better benchmarking of single cell RNA-seq protocols and analysis methods

Single cell RNA sequencing (scRNA-seq) technology has undergone rapid development in recent years, bringing with it new challenges in data processing and analysis. This has led to an explosion of tailored analysis methods for scRNA-seq to address various biological questions. However, the current lack of gold-standard benchmarking datasets makes it difficult for researchers to evaluate the performance of the many methods. Here, we designed and carried out a realistic benchmark experiment that included mixtures of single cells or pseudo-cells created by sampling admixtures of cells or RNA from 3 distinct cancer cell lines. Altogether we generated 10 datasets using a combination of droplet and plate-based scRNA-seq protocols, with varying data quality, population heterogeneity and noise levels. Using these benchmark datasets, we compared different protocols, evaluated the spike-in standard and multiple data analysis methods for tasks ranging from normalization and imputation, to clustering, trajectory analysis and data integration. Evaluation of methods across multiple datasets revealed some that performed well in general and others that suited specific situations. Our dataset and analysis provide a comprehensive comparison framework for benchmarking most popular scRNA-seq analysis tasks.

bioinformatics

Serum glycoprotein biomarker validation for esophageal adenocarcinoma and application to Barrett’s surveillance

BACKGROUND & AIMSEsophageal adenocarcinoma (EAC) is thought to develop from asymptomatic Barretts esophagus (BE) with a low annual rate of conversion. Current endoscopy surveillance for BE patients is probably not cost-effective. Previously, we discovered serum glycoprotein biomarker candidates which could discriminate BE patients from EAC. Here, we aimed to validate candidate serum glycoprotein biomarkers in independent cohorts, and to develop a biomarker panel for BE surveillance.\n\nMETHODSSerum glycoprotein biomarker candidates were measured in 301 serum samples collected from Australia (4 states) and USA (1 clinic) using lectin magnetic bead array (LeMBA) coupled multiple reaction monitoring mass spectrometry (MRM-MS). The area under receiver operating characteristic curve was calculated as a measure of discrimination, and multivariate recursive partitioning was used to formulate a multi-marker panel for BE surveillance.\n\nRESULTSDifferent glycoforms of complement C9 (C9), gelsolin (GSN), serum paraoxonase/arylesterase 1 (PON1) and serum paraoxonase/lactonase 3 (PON3) were validated as diagnostic glycoprotein biomarker candidates for EAC across both cohorts. A panel of 10 serum glycoproteins accurately discriminated BE patients not requiring intervention [BE+/-low grade dysplasia] from those requiring intervention [BE with high grade dysplasia (BE-HGD) or EAC]. Tissue expression of C9 was found to be induced in BE, dysplastic BE and EAC. In longitudinal samples from subjects that have progressed towards EAC, levels of serum C9 glycoforms were increased with disease progression.\n\nCONCLUSIONSFurther prospective clinical validation of the confirmed biomarker candidates in a large cohort is warranted. A first-line BE surveillance blood test may be developed based on these findings.\n\nAbbreviations

biochemistry

mixOmics: an R package for ‘omics feature selection and multiple data integration

The advent of high throughput technologies has led to a wealth of publicly available omics data coming from different sources, such as transcriptomics, proteomics, metabolomics. Combining such large-scale biological data sets can lead to the discovery of important biological insights, provided that relevant information can be extracted in a holistic manner. Current statistical approaches have been focusing on identifying small subsets of molecules (a molecular signature) to explain or predict biological conditions, but mainly for a single type of omics. In addition, commonly used methods are univariate and consider each biological feature independently.\n\nWe introduce mixOmics, an R package dedicated to the multivariate analysis of biological data sets with a specific focus on data exploration, dimension reduction and visualisation. By adopting a system biology approach, the toolkit provides a wide range of methods that statistically integrate several data sets at once to probe relationships between heterogeneous omics data sets. Our recent methods extend Projection to Latent Structure (PLS) models for discriminant analysis, for data integration across multiple omics data or across independent studies, and for the identification of molecular signatures. We illustrate our latest mixOmics integrative frameworks for the multivariate analyses of omics data available from the package.

bioinformatics