Search bioRxivSearch

Biology subjects

Afsari, B.

Publications and source records attributed to Afsari, B..

4 recordsLinked to original sources

REVA: a rank-based multi-dimensional measure of correlation

The neighbors principle implicit in any machine learning algorithm says that samples with similar labels should be close to one another in feature space as well. For example, while tumors are heterogeneous, tumors that have similar genomics profiles can also be expected to have similar responses to a specific therapy. Simple correlation coefficients provide an effective way to determine whether this principle holds when features and labels are both scalar, but not when either is multivariate. A new class of generalized correlation coefficients based on inter-point distances addresses this need and is called \"distance correlation\". There is only one rank-based distance correlation test available to date, and it is asymmetric in the samples, requiring that one sample be distinguished as a fixed point of reference. Therefore, we introduce a novel, nonparametric statistic, REVA, inspired by the Kendall rank correlation coefficient. We use U-statistic theory to derive the asymptotic distribution of the new correlation coefficient, developing additional large and finite sample properties along the way. To establish the admissibility of the REVA statistic, and explore the utility and limitations of our model, we compared it to the most widely used distance based correlation coefficient in a range of simulated conditions, demonstrating that REVA does not depend on an assumption of linearity, and is robust to high levels of noise, high dimensions, and the presence of outliers. We also present an application to real data, applying REVA to determine whether cancer cells with similar genetic profiles also respond similarly to a targeted therapeutic.\n\nAuthor summarySometimes a simple question arises: how does the distance between two samples in multivariate space compare to another scalar value associated with each sample. Here, we propose theory for a nonparametric test to statistically test this association. This test is independent of the scale of the scalar data, and thus generalizable to any comparison of samples with both high-dimensional data and a scalar. We apply the resulting statistic, REVA, to problems in cancer biology motivated by the model that cancer cells with more similar gene expression profiles to one another can be expected to have a more similar response to therapy.

bioinformatics

Non-invasive detection of upper tract urothelial carcinomas through the analysis of driver gene mutations and aneuploidy in urine

Upper tract urothelial carcinomas (UTUC) of the renal pelvis or ureter can be difficult to detect and challenging to diagnose. Here, we report the development and application of a non-invasive test for UTUC based on molecular analyses of DNA recovered from cells shed into the urine. The test, called UroSEEK, incorporates assays for mutations in eleven genes frequently mutated in urologic malignancies and for allelic imbalances on 39 chromosome arms. At least one genetic abnormality was detected in 75% of urinary cell samples from 56 UTUC patients but in only 0.5% of 188 samples from healthy individuals. The assay was considerably more sensitive than urine cytology, the current standard-of-care. UroSEEK therefore has the potential to be used for screening or to aid in diagnosis in patients at increased risk for UTUC, such as those exposed to herbal remedies containing the carcinogen aristolochic acid.

cancer biology

Non-invasive detection of bladder cancer through the analysis of driver gene mutations and aneuploidy

Current non-invasive approaches for bladder cancer (BC) detection are suboptimal. We report the development of non-invasive molecular test for BC using DNA recovered from cells shed into urine. This \"UroSEEK\" test incorporates assays for mutations in 11 genes and copy number changes on 39 chromosome arms. We first evaluated 570 urine samples from patients at risk for BC (microscopic hematuria or dysuria). UroSEEK was positive in 83% of patients that developed BC, but in only 7% of patients who did not develop BC. Combined with cytology, 95% of patients that developed BC were positive. We then evaluated 322 urine samples from patients soon after their BCs had been surgically resected. UroSEEK detected abnormalities in 66% of the urine samples from these patients, sometimes up to 4 years prior to clinical evidence of residual neoplasia, while cytology was positive in only 25% of such urine samples. The advantages of UroSEEK over cytology were particularly evident in low-grade tumors, wherein cytology detected none while UroSEEK detected 67% of 49 cases. These results establish the foundation for a new, non-invasive approach to the detection of BC in patients at risk for initial or recurrent disease.

cancer biology

Splice Expression Variation Analysis (SEVA) for Differential Gene Isoform Usage in Cancer

MotivationCurrent bioinformatics methods to detect changes in gene isoform usage in distinct phenotypes compare the relative expected isoform usage in phenotypes. These statistics model differences in isoform usage in normal tissues, which have stable regulation of gene splicing. Pathological conditions, such as cancer, can have broken regulation of splicing that increases the heterogeneity of the expression of splice variants. Inferring events with such differential heterogeneity in gene isoform usage requires new statistical approaches.\n\nResultsWe introduce Splice Expression Variability Analysis (SEVA) to model increased heterogeneity of splice variant usage between conditions (e.g., tumor and normal samples). SEVA uses a rank-based multivariate statistic that compares the variability of junction expression profiles within one condition to the variability within another. Simulated data show that SEVA is unique in modeling heterogeneity of gene isoform usage, and benchmark SEVAs performance against EBSeq, DiffSplice, and rMATS that model differential isoform usage instead of heterogeneity. We confirm the accuracy of SEVAin identifying known splice variants in head and neck cancer and perform cross-study validation of novel splice variants. A novel comparison of splice variant heterogeneity between subtypes of head and neck cancer demonstrated unanticipated similarity between the heterogeneity of gene isoform usage in HPV-positive and HPV-negative subtypes and anticipated increased heterogeneity among HPV-negative samples with mutations in genes that regulate the splice variant machinery.\n\nConclusionThese results show that SEVA accurately models differential heterogeneity of gene isoform usage from RNA-seq data.\n\nAvailabilitySEVA is implemented in the R/Bioconductor package GSReg.\n\nContactbahman@jhu.edu, favorov@sensi.org, ejfertig@jhmi.edu

genomics