Search bioRxivSearch

Biology subjects

Koo, P. K.

Publications and source records attributed to Koo, P. K..

4 recordsLinked to original sources

An Empirical Demonstration of Unsupervised Machine Learning in Species Delimitation

One major challenge to delimiting species with genetic data is successfully differentiating species divergences from population structure, with some current methods biased towards overestimating species numbers. Many fields of science are now utilizing machine learning (ML) approaches, and in systematics and evolutionary biology, supervised ML algorithms have recently been incorporated to infer species boundaries. However, these methods require the creation of training data with associated labels. Unsupervised ML, on the other hand, uses the inherent structure in data and hence does not require any user-specified training labels, thus providing a more objective approach to species delimitation. In the context of integrative taxonomy, we demonstrate the utility of three unsupervised ML approaches, specifically random forests, variational autoencoders, and t-distributed stochastic neighbor embedding, for species delimitation utilizing a short-range endemic harvestman taxon (Laniatores, Metanonychus). First, we combine mitochondrial data with examination of male genitalic morphology to identify a priori species hypotheses. Then we use single nucleotide polymorphism data derived from sequence capture of ultraconserved elements (UCEs) to test the efficacy of unsupervised ML algorithms in successfully identifying a priori species, comparing results to commonly used genetic approaches. Finally, we use two validation methods to assess a priori species hypotheses using UCE data. We find that unsupervised ML approaches successfully cluster samples according to species level divergences and not to high levels of population structure, while standard model-based validation methods over-split species, in some instances suggesting that all sampled individuals are distinct species. Moreover, unsupervised ML approaches offer the benefits of better data visualization in two-dimensional space and the ability to accommodate various data types. We argue that ML methods may be better suited for species delimitation relative to currently used model-based validation methods, and that species delimitation in a truly integrative framework provides more robust final species hypotheses relative to separating delimitation into distinct \"discovery\" and \"validation\" phases. Unsupervised ML is a powerful analytical approach that can be incorporated into many aspects of systematic biology, including species delimitation. Based on results of our empirical dataset, we make several taxonomic changes including description of a new species.

evolutionary biology

Inferring Sequence-Structure Preferences of RNA-Binding Proteins with Convolutional Residual Networks

To infer the sequence and RNA structure specificities of RNA-binding proteins (RBPs) from experiments that enrich for bound sequences, we introduce a convolutional residual network which we call ResidualBind. ResidualBind significantly outperforms previous methods on experimental data from many RBP families. We interrogate ResidualBind to identify what features it has learned from high-affinity sequences with saliency analysis along with 1st-order and 2nd-order in silico mutagenesis. We show that in addition to sequence motifs, ResidualBind learns a model that includes the number of motifs, their spacing, and both positive and negative effects of RNA structure context. Strikingly, ResidualBind learns RNA structure context, including detailed base-pairing relationships, directly from sequence data, which we confirm on synthetic data. ResidualBind is a powerful, flexible, and interpretable model that can uncover cis-recognition preferences across a broad spectrum of RBPs.

genomics

Representation Learning of Genomic Sequence Motifs with Convolutional Neural Networks

Although convolutional neural networks (CNNs) have been applied to a variety of computational genomics problems, there remains a large gap in our understanding of how they build representations of regulatory genomic sequences. Here we perform systematic experiments on synthetic sequences to reveal how CNN architecture, specifically convolutional filter size and max-pooling, influences the extent that sequence motif representations are learned by first layer filters. We find that CNNs designed to foster hierarchical representation learning of sequence motifs - assembling partial features into whole features in deeper layers - tend to learn distributed representations, i.e. partial motifs. On the other hand, CNNs that are designed to limit the ability to hierarchically build sequence motif representations in deeper layers tend to learn more interpretable localist representations, i.e. whole motifs. We then validate that this representation learning principle established from synthetic sequences generalizes to in vivo sequences.

genomics

Inferring Functional Neural Connectivity With Deep Residual Convolutional Networks

Measuring synaptic connectivity in large neuronal populations remains a major goal of modern neuroscience. While this connectivity is traditionally revealed by anatomical methods such as electron microscopy, an efficient alternative is to computationally infer functional connectivity from recordings of neural activity. However, these statistical techniques still require further refinement before they can be reliably applied to real data. Here, we report significant improvements to a deep learning method for functional connectomics, as assayed on synthetic ChaLearn Connectomics data. The method, which integrates recent advances in convolutional neural network architecture and model-free partial correlation coefficients, outperforms published methods on competition data and can achieve over 90% precision at 1% recall on validation datasets. This suggests that future application of the model to in vivo whole-brain imaging data in larval zebrafish could reliably recover on the order of 106 synaptic connections with a 10% false discovery rate. The model also generalizes to networks with different underlying connection probabilities and should scale well when parallelized across multiple GPUs. The method offers real potential as a statistical complement to existing experiments and circuit hypotheses in neuroscience.

neuroscience