Search bioRxivSearch

Biology subjects

Miecznikowski, J. C.

Publications and source records attributed to Miecznikowski, J. C..

3 recordsLinked to original sources

MOSCATO: A Supervised Approach for Analyzing Multi-Omic Single-Cell Data

Advancements in genomic sequencing continually improve personalized medicine in complex diseases. Recent breakthroughs generate multiple types of signatures (or multi-omics) from each cell, producing different data omic types per single-cell experiment. We introduce MOSCATO, a technique for selecting features across multi-omic single-cell datasets that relate to clinical outcomes. For example, we leverage penalization concepts often used in multi-omic network analytics to accommodate the high-dimensionality where multiple-testing is likely underpowered. We organize the data into multi-dimensional tensors where the dimensions correspond to the different omic types. Using the outcome and the single-cell tensors, we perform regularized tensor regression to return a variable set for each omic type that forms the clinically-associated network. Robustness is assessed over simulations based on available single-cell simulation methods. Real data comparing healthy subjects versus subjects with leukemia is also considered in order to identify genes associated with the disease. The flexibility of our approach enables future extensions on distributional assumptions and covariate adjustments. This algorithm may identify clinically-relevant genetic patterns on a cellular-level that span multiple layers of sequencing data and ultimately inform highly precise therapeutic targets in complex diseases. Code to perform MOSCATO and replicate the real data application is publicly available on GitHub at https://github.com/lorinmil/MOSCATO and https://github.com/lorinmil/MOSCATOLeukemiaExample.

bioinformatics

Balanced Functional Module Detection in Genomic Data

High dimensional genomic data can be analyzed to understand the effects of multiple variables on a target variable such as a clinical outcome, risk factor or diagnosis. Of special interest are functional modules, cooperating sets of variables affecting the target. Graphical models of various types are often useful for characterizing such networks of variables. In other applications such as social networks, the concept of balance in undirected signed graphs characterizes the consistency of associations within the network. To extend this concept to applications where a set of predictor variables influences an outcome variable, we define balance for functional modules. This property specifies that the module variables have a joint effect on the target outcome with no internal conflict, an efficiency that evolution may use for selection in biological networks. We show that for this class of graphs, observed correlations directly reflect paths in the underlying graph. Consequences of the balance property are exploited to implement a new module discovery algorithm, bFMD, which selects a subset of variables from highdimensional data that compose a balanced functional module. Our bFMD algorithm performed favorably in simulations as compared to other module detection methods that do not consider balance properties. Additionally, bFMD detected interpretable results in a real application for RNA-seq data obtained from The Cancer Genome Atlas (TCGA) for Uterine Corpus Endometrial Carcinoma using the percentage of tumor invasion as the target outcome of interest. bFMD detects sparse sets of variables within highdimensional datasets such that interpretability may be favorable as compared to other similar methods by leveraging balance properties used in other graphical applications.

bioinformatics

Filtering Variables for Supervised Sparse Network Analysis

MotivationWe present a method for dimension reduction designed to filter variables or features such as genes considered to be irrelevant for a downstream analysis designed to detect supervised gene networks in sparse settings. This approach can improve interpret-ability for a variety of analysis methods. We present a method to filter genes and transcripts prior to network analysis. This method has applications in a setting where the downstream analysis may include sparse canonical correlation analysis. ResultsFiltering methods specifically for cluster and network analysis are introduced and compared by simulating modular networks with known statistical properties. Our proposed method performs favorably eliminating irrelevant features but maintaining important biological signal under a variety of different signal settings. We show that the speed and accuracy of methods such as sparse canonical correlation are increased after filtering, thus greatly improving the scalability of these approaches. AvailabilityCode for performing the gene filtering algorithm described in this manuscript may be accessed through the geneFiltering R package available on Github at https://github.com/lorinmil/geneFiltering. Functions are available to filter genes and perform simulations of a network system. For access to the data used in this manuscript, contact corresponding author. Contactlorinmil@buffalo.edu, jcm38@buffalo.edu, fzhang8@buffalo.edu, and dlt6@buffalo.edu

bioinformatics