Search bioRxivSearch

Biology subjects

Song, J. S.

Publications and source records attributed to Song, J. S..

6 recordsLinked to original sources

SequencEnG: an Interactive Knowledge Base of Sequencing Techniques

Next-generation sequencing (NGS) techniques are revolutionizing biomedical research by providing powerful methods for generating genomic and epigenomic profiles. The rapid progress is posing an acute challenge to students and researchers to stay acquainted with the numerous available methods. We have developed an interactive online educational resource called SequencEnG (acronym for Sequencing Techniques Engine for Genomics) to provide a tree-structured knowledge base of 66 different sequencing techniques and step-by-step NGS data analysis pipelines comparing popular tools. SequencEnG is designed to facilitate barrier-free learning of current NGS techniques and provides a user-friendly interface for searching through experimental and analysis methods. SequencEnG is part of the project KnowEnG (Knowledge Engine for Genomics) and is freely available at http://education.knoweng.org/sequenceng/.

bioinformatics

A unified computational framework for modeling genome-wide nucleosome landscape

Nucleosomes form the fundamental building blocks of eukaryotic chromatin, and previous attempts to understand the principles governing their genome-wide distribution have spurred much interest and debate in biology. In particular, the precise role of DNA sequence in shaping local chromatin structure has been controversial. This paper rigorously quantifies of the contribution of hitherto-debated sequence features - including G+C content, 10.5-bp periodicity, and poly(dA:dT) tracts - to three distinct aspects of genome-wide nucleosome landscape: occupancy, translational positioning and rotational positioning. Our computational framework simultaneously learns nucleosome number and nucleosome-positioning energy from genome-wide nucleosome maps. In contrast to other previous studies, our model can predict both in-vitro and in-vivo nucleosome maps in S. cerevisiae. We find that although G+C content is the primary determinant of MNase-derived nucleosome occupancy, MNase digestion biases may substantially influence this GC dependence. By contrast, poly(dA:dT) tracts are seen to deter nucleosome formation, regardless of the experimental method used. We further show that the 10.5-bp nucleotide periodicity facilitates rotational but not translational positioning. Applying our method to in-vivo nucleosome maps demonstrates that, for a subset of genes, the regularly-spaced nucleosome arrays observed around transcription start sites can be partially recapitulated by DNA sequence alone. Finally, in-vivo nucleosome occupancy derived from MNase-seq experiments around transcription termination sites can be mostly explained by the genomic sequence. Implications of these results and potential extensions of the proposed computational framework are discussed

bioinformatics

High accuracy label-free classification of kinetic cell states from holographic cytometry

Digital holographic microscopy permits live and label-free visualization of adherent cells. Here we report the application of this approach for high accuracy kinetic quantitative cytometry. We identify twenty-six label-free optical and morphological features that are biologically independent. When used as a basis for machine learning, these features allow blind single cell classification with up to 95% accuracy. We present methods to control for inherent holographic noise, thereby establishing a set of reliable quantitative features. Together, these contributions permit continuous digital holographic cytometry for three or more days. Applying our approach to human melanoma cells treated with a panel of cancer therapeutics, we can track the response of each cell, simultaneously classifying multiple behaviors such as cell cycle length, motility, apoptosis, senescence, and heterogeneity of response to each therapeutic. Importantly, we demonstrate relationships between these phenotypes over time. This work thus provides an experimental and computational roadmap for low cost live-cell imaging and kinetic classification of heterogeneous adherent cell populations.

cancer biology

ClusterEnG: An interactive educational web resource for clustering big data

SummaryClustering is one of the most common techniques used in data analysis to discover hidden structures by grouping together data points that are similar in some measure into clusters. Although there are many programs available for performing clustering, a single web resource that provides both state-of-the-art clustering methods and interactive visualizations is lacking. ClusterEnG (acronym for Clustering Engine for Genomics) provides an interface for clustering big data and interactive visualizations including 3D views, cluster selection and zoom features. ClusterEnG also aims at educating the user about the similarities and differences between various clustering algorithms and provides clustering tutorials that demonstrate potential pitfalls of each algorithm. The web resource will be particularly useful to scientists who are not conversant with computing but want to understand the structure of their data in an intuitive manner.\n\nAvailabilityClusterEnG is part of a bigger project called KnowEnG (Knowledge Engine for Genomics) and is available at http://education.knoweng.org/clustereng.\n\nContactsongi@illinois.edu

bioinformatics

Emergent community agglomeration from data set geometry

In the statistical learning language, samples are snapshots of random vectors drawn from some unknown distribution. Such vectors usually reside in a high-dimensional Euclidean space, and thus, the \"curse of dimensionality\" often undermines the power of learning methods, including community detection and clustering algorithms, that rely on Euclidean geometry. This paper presents the idea of effective dissimilarity transformation (EDT) on empirical dissimilarity hyperspheres and studies its effects using synthetic and gene expression data sets. Iterating the EDT turns a static data distribution into a dynamical process purely driven by the empirical data set geometry and adaptively ameliorates the curse of dimensionality, partly through changing the topology of a Euclidean feature space [R]n into a compact hypersphere Sn. The EDT often improves the performance of hierarchical clustering via the automatic grouping information emerging from global interactions of data points. The EDT is not restricted to hierarchical clustering, and other learning methods based on pairwise dissimilarity should also benefit from the many desirable properties of EDT.\n\nPACS numbers: 89.20.Ff, 87.85.mg

bioinformatics

Maximum Entropy Methods for Extracting the Learned Features of Deep Neural Networks

New architectures of multilayer artificial neural networks and new methods for training them are rapidly revolutionizing the application of machine learning in diverse fields, including business, social science, physical sciences, and biology. Interpreting deep neural networks, however, currently remains elusive, and a critical challenge lies in understanding which meaningful features a network is actually learning. We present a general method for interpreting deep neural networks and extracting network-learned features from input data. We describe our algorithm in the context of biological sequence analysis. Our approach, based on ideas from statistical physics, samples from the maximum entropy distribution over possible sequences, anchored at an input sequence and subject to constraints implied by the empirical function learned by a network. Using our framework, we demonstrate that local transcription factor binding motifs can be identified from a network trained on ChIP-seq data and that nucleosome positioning signals are indeed learned by a network trained on chemical cleavage nucleosome maps. Imposing a further constraint on the maximum entropy distribution also allows us to probe whether a network is learning global sequence features, such as the high GC content in nucleosome-rich regions. This work thus provides valuable mathematical tools for interpreting and extracting learned features from feed-forward neural networks.

bioinformatics