Search bioRxiv⌕ Search

Biology subjects

Ortiz-Velez, A. N.

Publications and source records attributed to Ortiz-Velez, A. N..

2 recordsLinked to original sources

Data-Driven Mathematical Approach for Removing Rare Features in Zero-Inflated Datasets

Sparse feature tables, in which many features are present in very few samples, are common in big biological data (e.g., metagenomics, transcriptomics). Ignoring the problem of zero-inflation can result in biased statistical estimates and decrease power in downstream analyses. Zeros are also a particular issue for compositional data analysis using log-ratios since the log of zero is undefined. Researchers typically deal with zero-inflated data by removing low frequency features, but the thresholds for removal differ markedly between studies with little or no justification. Here, we present CurvCut, a data-driven mathematical approach to zero-inflated feature removal based on curvature analysis of a "ball rolling down a hill", where the hill is a histogram of feature distribution. These histograms typically contain a point of regime change, a discontinuity with a sharp change in the characteristics of the distribution, that can be used as a cutoff point for low frequency feature removal that considers the data-specific nature of the feature distribution. Our results show that CurvCut works well across a variety of biological data types, including ones with both right- and left-skewed feature distributions, and rapidly generates clear visual results allowing researchers to select data-appropriate cutoffs for feature removal.

bioinformatics↗

AutoPhy: Automated phylogenetic identification of novel protein subfamilies

Phylogenetic analysis of protein sequences provides a powerful means of identifying novel protein functions and subfamilies, and for identifying and resolving annotation errors. However, automation of functional clustering based on phylogenetic trees has been challenging, and most of it is done manually. Clustering phylogenetic trees usually requires the delineation of tree-based thresholds (e.g., distances), leading to an ad hoc problem. We propose a new phylogenetic clustering approach that identifies clusters without using ad hoc distances or other pre-defined values. Our workflow combines uniform manifold approximation and projection (UMAP) with Gaussian mixture models as a k-means like procedure to automatically group sequences into clusters. We then apply a "second pass" clade identification algorithm to resolve non-monophyletic groups. We tested our approach with several well-curated protein families (outer membrane porins, acyltransferase, and dehydrogenases) and showed our automated methods recapitulated known subfamilies. We also applied our methods to a broad range of different protein families from multiple databases, including Pfam, PANTHER, and UNIPROT. Our results showed that AutoPhy rapidly generated monophyletic clusters (subfamilies) within phylogenetic trees evolving at very different rates both within and among phylogenies. The phylogenetic clusters generated by AutoPhy resolved misannotations, determined new protein functional groups, and detected novel viral strains.

bioinformatics↗