Search bioRxivSearch

Biology subjects

Castillo, S.

Publications and source records attributed to Castillo, S..

2 recordsLinked to original sources

An Empirical Demonstration of Unsupervised Machine Learning in Species Delimitation

One major challenge to delimiting species with genetic data is successfully differentiating species divergences from population structure, with some current methods biased towards overestimating species numbers. Many fields of science are now utilizing machine learning (ML) approaches, and in systematics and evolutionary biology, supervised ML algorithms have recently been incorporated to infer species boundaries. However, these methods require the creation of training data with associated labels. Unsupervised ML, on the other hand, uses the inherent structure in data and hence does not require any user-specified training labels, thus providing a more objective approach to species delimitation. In the context of integrative taxonomy, we demonstrate the utility of three unsupervised ML approaches, specifically random forests, variational autoencoders, and t-distributed stochastic neighbor embedding, for species delimitation utilizing a short-range endemic harvestman taxon (Laniatores, Metanonychus). First, we combine mitochondrial data with examination of male genitalic morphology to identify a priori species hypotheses. Then we use single nucleotide polymorphism data derived from sequence capture of ultraconserved elements (UCEs) to test the efficacy of unsupervised ML algorithms in successfully identifying a priori species, comparing results to commonly used genetic approaches. Finally, we use two validation methods to assess a priori species hypotheses using UCE data. We find that unsupervised ML approaches successfully cluster samples according to species level divergences and not to high levels of population structure, while standard model-based validation methods over-split species, in some instances suggesting that all sampled individuals are distinct species. Moreover, unsupervised ML approaches offer the benefits of better data visualization in two-dimensional space and the ability to accommodate various data types. We argue that ML methods may be better suited for species delimitation relative to currently used model-based validation methods, and that species delimitation in a truly integrative framework provides more robust final species hypotheses relative to separating delimitation into distinct \"discovery\" and \"validation\" phases. Unsupervised ML is a powerful analytical approach that can be incorporated into many aspects of systematic biology, including species delimitation. Based on results of our empirical dataset, we make several taxonomic changes including description of a new species.

evolutionary biology

Principal Metabolic Flux Mode Analysis

MotivationIn the analysis of metabolism using omics data, two distinct and complementary approaches are frequently used: Principal component analysis (PCA) and Stoichiometric flux analysis. PCA is able to capture the main modes of variability in a set of experiments and does not make many prior assumptions about the data, but does not inherently take into account the flux mode structure of metabolism. Stoichiometric flux analysis methods, such as Flux Balance Analysis (FBA) and Elementary Mode Analysis, on the other hand, produce results that are readily interpretable in terms of metabolic flux modes, however, they are not best suited for exploratory analysis on a large set of samples.\n\nResultsWe propose a new methodology for the analysis of metabolism, called Principal Metabolic Flux Mode Analysis (PMFA), which marries the PCA and Stoichiometric flux analysis approaches in an elegant regularized optimization framework. In short, the method incorporates a variance maximization objective form PCA coupled with a Stoichiometric regularizer, which penalizes projections that are far from any flux modes of the network. For interpretability, we also introduce a sparse variant of PMFA that favours flux modes that contain a small number of reactions. Our experiments demonstrate the versatility and capabilities of our methodology.\n\nAvailabilityMatlab software for PMFA and SPMFA is available in https://github.com/ aalto-ics-kepaco/PMFA.\n\nContactsahely@iitpkd.ac.in, juho.rousu@aalto.fi, Peter.Blomberg@vtt.fi, Sandra.Castillo@vtt.fi\n\nSupplementary informationDetailed results are in Supplementary files. Supplementary data are available at https://github.com/aalto-ics-kepaco/PMFA/blob/master/Results.zip.

bioinformatics