Search bioRxivSearch

Biology subjects

El-Manzalawy, Y.

Publications and source records attributed to El-Manzalawy, Y..

2 recordsLinked to original sources

Proxi: a Python package for proximity network inference from metagenomic data

Summary: Recent technological advances in high-throughput metagenomic sequencing have provided unique opportunities for studying the diversity and dynamics of microbial communities under different health or environmental conditions. Graph-based representation of metagenomic data is a promising direction not only for analyzing microbial interactions but also for a broad range of machine learning tasks including feature selection, classification, clustering, anomaly detection, and dimensionality reduction. We present Proxi, an open source Python package for learning different types of proximity graphs from metagenomic data. Currently, three types of proximity graphs are supported: k-nearest neighbor (k-NN) graphs; radius-nearest neighbor (r-NN) graphs; and perturbed k-nearest neighbor (pk-NN) graphs.\n\nAvailability: Proxi Python source code is freely available at https://bitbucket.org/idsrlab/proxi/.\n\nContact: yme2@psu.edu\n\nSupplementary information: Tutorials and online documentation are available at https://proxi.readthedocs.io

systems biology

Min-Redundancy and Max-Relevance Multi-view Feature Selection for Predicting Ovarian Cancer Survival using Multi-omics Data

BackgroundLarge-scale collaborative precision medicine initiatives (e.g., The Cancer Genome Atlas (TCGA)) are yielding rich multi-omics data. Integrative analyses of the resulting multi-omics data, such as somatic mutation, copy number alteration (CNA), DNA methylation, miRNA, gene expression, and protein expression, offer the tantalizing possibilities of realizing the potential of precision medicine in cancer prevention, diagnosis, and treatment by substantially improving our understanding of underlying mechanisms as well as the discovery of novel biomarkers for different types of cancers. However, such analyses present a number of challenges, including the heterogeneity of data types, and the extreme high-dimensionality of omics data.\n\nMethodsIn this study, we propose a novel framework for integrating multi-omics data based on multi-view feature selection, an emerging research problem in machine learning research. We also present a novel multi-view feature selection algorithm, MRMR-mv, which adapts the well-known Min-Redundancy and Maximum-Relevance (MRMR) single-view feature selection algorithm for the multi-view settings.\n\nResultsWe report results of experiments on the task of building a predictive model of cancer survival from an ovarian cancer multi-omics dataset derived from the TCGA database. Our results suggest that multi-view models for predicting ovarian cancer survival outperform both view-specific models (i.e., models trained and tested using one multi-omics data source) and models based on two baseline data fusion methods.\n\nConclusionsOur results demonstrate the potential of multi-view feature selection in integrative analyses and predictive modeling from multi-omics data.

bioinformatics