Search bioRxivSearch

Biology subjects

Varma, S.

Publications and source records attributed to Varma, S..

3 recordsLinked to original sources

Integrative analysis of pharmacogenomics in major cancer cell line databases using CellMinerCDB

As precision medicine demands molecular determinants of drug response, CellMinerCDB provides (https://discover.nci.nih.gov/cellminercdb/) a web-based portal for multiple forms of pharmacological, molecular, and genomic analyses, unifying the richest cancer cell line datasets (NCI-60, NCI-SCLC, Sanger/MGH GDSC, and Broad CCLE/CTRP). CellMinerCDB enables genomic and pharmacological data queries for identifying pharmacogenomic determinants, drug signatures, and gene regulatory networks for researchers without requiring specialized bioinformatics support. It leverages overlaps of cell lines and tested drugs to allow assessment of data reproducibility. It builds on the complementarity and strength of each dataset. A panel of 41 drugs evaluated in parallel in the NCI-60 and GDSC is reported, supporting drug reproducibility across databases, repositioning of bisacodyl and acetalax for triple negative breast cancer, and identifying novel drug response determinants and genomic signatures for topoisomerase inhibitors and schweinfurthins in development. CellMinerCDB also allowed the identification of LIX1L as a novel mesenchymal gene regulating cellular migration and invasiveness.

bioinformatics

Rapid proteotyping reveals cancer biology and drug response determinants in the NCI-60 cells

We describe the rapid and reproducible acquisition of quantitative proteome maps for the NCI-60 cancer cell lines and their use to reveal cancer biology and drug response determinants. Proteome datasets for the 60 cell lines were acquired in duplicate within 30 working days using pressure cycling technology and SWATH mass spectrometry. We consistently quantified 3,171 proteotypic proteins annotated in the SwissProt database across all cell lines, generating a data matrix with 0.1% missing values, allowing analyses of protein complexes and pathway activities across all the cancer cells. Systematic and integrative analysis of the genetic variation, mRNA expression and proteomic data of the NCI-60 cancer cell lines uncovered complementarity between different types of molecular data in the prediction of the response to 240 drugs. We additionally identified novel proteomic drug response determinants for clinically relevant chemotherapeutic and targeted therapies. We anticipate that this study represents a significant advance toward the translational application of proteotypes, which reveal biological insights that are easily missed in the absence of proteomic data.

systems biology

Blind estimation and correction of microarray batch effect

Microarray batch effect (BE) has been the primary bottleneck for large-scale integration of data from multiple experiments. Current BE correction methods either need known batch identities (ComBat)or have the potential to overcorrect, by removing true but unknown biological differences (SVA).Even though the effects of technical differences on measured expression have been published, there are no BE correction algorithms that take the approach of predicting technical effects from parameters computed from a fixed reference sample set. We show that a set of signatures, each of which is a vector the length of the number of probes, calculated on a Reference set of microarray samples can predict much of the batch effect in other Validation sets. We present a rationale of selecting a Reference set of samples designed to estimate technical differences without removing biological differences. Putting both together, we introduce the Batch Effect Signature Correction (BESC) algorithm that uses the BES calculated on the Reference set to efficiently predict and remove BE. Using two independent Validation sets, we show that BESC is capable of removing batch effect without removing unknown but true biological differences. Much of the variations due to batch effect is shared between different microarray datasets. That shared information can be used to predict signatures (i.e. directions of perturbation) due to batch effect in new datasets. The correction is blind (without needing to re-compute the parameters on new samples to be corrected), single sample, (each sample is corrected independently of each other) and conservative (only those perturbations known to be likely to be due to technical differences are removed ensuring that unknown but important biological differences are maintained). Those three characteristics make it ideal for high-throughput correction of samples for a microarray data repository. An R Package besc implementing the algorithm is available from http://explainbio.com.

bioinformatics