Search bioRxivSearch

Biology subjects

Peng, J.

Publications and source records attributed to Peng, J..

At least 19 recordsLinked to original sources

Identifying Emerging Phenomenon in Plant Long Temporal Phenotyping Experiments

The rapid improvement of phenotyping capability, accuracy, and throughput have greatly increased the volume and diversity of phenomics data. A remaining challenge is an efficient way to identify phenotypic patterns to improve our understanding of the quantitative variation of complex phenotypes, and to attribute gene functions. To address this challenge, we developed a new algorithm to identify emerging phenomena from large-scale temporal plant phenotyping experiments. An emerging phenomenon is defined as a group of genotypes who exhibit a coherent phenotype pattern during a relatively short time. Emerging phenomena are highly transient and diverse, and are dependent in complex ways on both environmental conditions and development. Identifying emerging phenomena may help biologists to examine potential relationships among phenotypes and genotypes in a genetically diverse population and to associate such relationships with the change of environments or development. We present an emerging phenomenon identification tool called Temporal Emerging Phenomenon Finder (TEP-Finder). Using large-scale longitudinal phenomics data as input, TEP-Finder first encodes the complicated phenotypic patterns into a dynamic phenotype network. Then, emerging phenomena in different temporal scales are identified from dynamic phenotype network using a maximal clique based approach. Meanwhile, a directed acyclic network of emerging phenomena is composed to model the relationships among the emerging phenomena. The experiment that compares TEP-Finder with two state-of-art algorithms shows that the emerging phenomena identified by TEP-Finder are more functionally specific, robust, and biologically significant. The source code, manual, and sample data of TEP-Finder are all available at: http://phenomics.uky.edu/TEP-Finder/.

bioinformatics

Combining Gene Ontology with Deep Neural Networks to Enhance the Clustering of Single Cell RNA-Seq Data

BackgroundSingle cell RNA sequencing (scRNA-seq) is applied to assay the individual transcriptomes of large numbers of cells. The gene expression at single-cell level provides an opportunity for better understanding of cell function and new discoveries in biomedical areas. To ensure that the single-cell based gene expression data are interpreted appropriately, it is crucial to develop new computational methods.\n\nResultsIn this article, we try to construct the structure of neural networks based on the prior knowledge of Gene Ontology (GO). By integrating GO with both unsupervised and supervised models, two novel methods are proposed, named GOAE (Gene Ontology AutoEncoder) and GONN (Gene Ontology Neural Network) respectively, for clustering of scRNA-seq data.\n\nConclusionsThe evaluation results show that the proposed models outperform some state-of-the-art approaches. Furthermore, incorporating with GO, we provide an opportunity to interpret the underlying biological mechanism behind the neural network-based model.

bioinformatics

Control of synaptic specificity by limiting promiscuous synapse formation

The ability of neurons to distinguish appropriate from inappropriate synaptic partners in their local environment is fundamental to the proper assembly and function of neural circuits. How synaptic partner selection is regulated is a longstanding question in Neurobiology. A prevailing hypothesis is that appropriate partners express complementary molecules that match them together and promote synaptogenesis. Dpr and DIP IgSF proteins bind heterophilically and are expressed in a complementary manner between synaptic partners in the Drosophila visual system. Here, we show that in the lamina, DIP mis-expression is sufficient to promote synapse formation with Dpr-expressing neurons, and that DIP proteins are not necessary for synaptogenesis but rather function to prevent ectopic synapse formation. These findings indicate that Dpr-DIP interactions regulate synaptic specificity by biasing synapse formation towards specific cell-types. We propose that synaptogenesis occurs independent of synaptic partner choice, and that precise synaptic connectivity is established by limiting promiscuous synapse formation.

neuroscience

Identifying Representative Network Motifs for Inferring Higher-order Structure of Biological Networks

Network motifs are recurring significant patterns of inter-connections, which are recognized as fundamental units to study the higher-order organizations of networks. However, the principle of selecting representative network motifs for local motif based clustering remains largely unexplored. We present a scalable algorithm called FSM for network motif discovery. FSM accelerates the motif discovery process by effectively reducing the number of times to do subgraph isomorphism labeling. Multiple heuristic optimizations for subgraph enumeration and subgraph classification are also adopted in FSM to further improve its performance. Experimental results show that FSM is more efficient than the compared models on computational efficiency and memory usage. Furthermore, our experiments indicate that large and frequent network motifs may be more appropriate to be selected as the representative network motifs for discovering higher-order organizational structures in biological networks than small or low-frequency network motifs.

bioinformatics

Dissecting differential signals in high-throughput data from complex tissues

Samples from clinical practices are often mixtures of different cell types. The high-throughput data obtained from these samples are thus mixed signals. The cell mixture brings complications to data analysis, and will lead to biased results if not properly accounted for. We develop a method to model the high-throughput data from mixed, heterogeneous samples, and to detect differential signals. Our method allows flexible statistical inference for detecting a variety of cell-type specific changes. Extensive simulation studies and analyses of two real datasets demonstrate the favorable performance of our proposed method compared with existing ones serving similar purpose.

bioinformatics

Flexibility and rigidity index for chromosome packing, flexibility and dynamics analysis

MotivationThe packing of genomic DNA from double string into highly-order hierarchial assemblies has great impact on chromosome flexibility, dynamics and functions. The open and accessible regions of chromosome are the primary binding positions for regulatory elements and are crucial to nuclear processes and biological functions.\n\nResultsMotivated by the success of flexibility-rigidity index (FRI) in biomolecular flexibility analysis and drug design, we propose a FRI based model for quantitatively characterizing the chromosome flexibility. Based on the Hi-C data, a flexibility index for each locus can be evaluated. Physically, the flexibility is tightly related to the packing density. Highly compacted regions are usually more rigid, while loosely packed regions are more flexible. Indeed, a strong correlation is found between our flexibility index and DNase and ATAC values, which are measurements for chromosome accessibility. Recently, Gaussian network model (GNM) is applied to analyze the chromosome accessibility and a mobility profile has been proposed to characterize the chromosome flexibility. Compared with GNM, our FRI is slightly more accurate (1% to 2% increase) and significantly more efficient in both computational time and costs. For a 5kb resolution Hi-C data, the flexibility evaluation process only takes FRI a few minutes on a single-core processor. In contrast, GNM requires 1.5 hours on 10 CPUs. Moreover, interchromosome information can be easily incorporated into the flexibility evaluation, thus further enhance the accuracy of our FRI. In contrast, the consideration of interchromosome information into GNM will significantly increase the size of its Laplacian matrix, thus computationally extremely challenging for the current GNM.\n\nAvailabilityThe software is available at https://github.com/jiajiepeng/FRI_chrFle.\n\nContactxiakelin@ntu.edu.sg; jiajiepeng@nwpu.edu.cn

bioinformatics

Zebrafish hhex null mutant develops an intrahepatic intestinal tube due to de-repression of cdx1b and pdx1

The hepatopancreatic duct (HPD) system links the liver and pancreas to the intestinal tube and is composed of the extrahepatic biliary duct, gallbladder and pancreatic duct. Haematopoietically-expressed-homeobox (Hhex) protein plays an essential role in the establishment of HPD, however, the molecular mechanism remains elusive. Here we show that zebrafish hhex-null mutants fail to develop the HPD system characterized by lacking the biliary marker Annexin A4 and the HPD marker sox9b. The mutant HPD system is replaced by an intrahepatic intestinal tube characterized by expressing the intestinal marker fatty-acid-binding-protein 2a (fabp2a). Cell lineage analysis showed that this intrahepatic intestinal tube is not originated from hepatocytes or cholangiocytes. Further analysis revealed that cdx1b and pdx1 were expressed ectopically in the intrahepatic intestinal tube and knockdown of cdx1b and pdx1 restored the expression of sox9b in the mutant. Chromatin-immunoprecipitation analysis shows that Hhex binds to the promoters of pdx1 and cdx1b genes to repress their expression. We therefore propose that Hhex, Cdx1b and Pdx1 form a genetic network governing the patterning and morphogenesis of the HPD and digestive tract systems in zebrafish.

developmental biology

Stratification of amyotrophic lateral sclerosis patients: a crowdsourcing approach

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease with substantial heterogeneity in clinical presentation with an urgent need for better stratification tools for clinical development and care. In this study we used a crowdsourcing approach to address the problem of ALS patient stratification. The DREAM Prize4Life ALS Stratification Challenge was a crowdsourcing initiative using data from >10,000 patients from completed ALS clinical trials and 1479 patients from community-based patient registers. Challenge participants used machine learning and clustering techniques to predict ALS progression and survival. By developing new approaches, the best performing teams were able to predict disease outcomes better than currently available methods. At the same time, the integration of clustering components across methods led to the emergence of distinct consensus clusters, separating patients into four consistent groups, each with its unique predictors for classification. This analysis reveals for the first time the potential of a crowdsourcing approach to uncover covert patient sub-populations, and to accelerate disease understanding and therapeutic development.

bioinformatics

Neural Data Visualization for Scalable and Generalizable Single Cell Analysis

Single-cell RNA sequencing is becoming effective and accessible as emerging technologies push its scale to millions of cells and beyond. Visualizing the landscape of single cell expression has been a fundamental tool in single cell analysis. However, standard methods for visualization, such as t-stochastic neighbor embedding (t-SNE), not only lack scalability to data sets with millions of cells, but also are unable to generalize to new cells, an important ability for transferring knowledge across fast-accumulating data sets. We introduce net-SNE, which trains a neural network to learn a high quality visualization of single cells that newly generalizes to unseen data. While matching the visualization quality of t-SNE on 14 benchmark data sets of varying sizes, from hundreds to 1.3 million cells, net-SNE also effectively positions previously unseen cells, even when an entire subtype is missing from the initial data set or when the new cells are from a different sequencing experiment. Furthermore, given a \"reference\" visualization, net-SNE can vastly reduce the computational burden of visualizing millions of single cells from multiple days to just a few minutes of runtime. Our work provides a general framework for newly bootstrapping single cell analysis from existing data sets.

bioinformatics

Deciphering signaling specificity with interpretable deep neural networks

Protein kinase phosphorylation is a prevalent post-translational modification (PTM) regulating protein function and transmitting signals throughout the cell. Defective signal transductions, which are associated with protein phosphorylation, have been revealed to link to many human diseases, such as cancer. Defining the organization of the phosphorylation-based signaling network and, in particular, identifying kinase-specific substrates can help reveal the molecular mechanism of the signaling network. Here, we present DeepSignal, a deep learning framework for predicting the substrate specificity for kinase/SH2 sequences with or without mutations. Empowered by the memory and selection mechanism of recurrent neural network, DeepSignal can identify important specificity-defining residues to predict kinase specificity and changes upon mutations. Evaluated on several public benchmark datasets, DeepSignal significantly outperforms current methods on predicting substrate specificity on both kinase and SH2 domains. Further analysis in The Cancer Genome Atlas (TCGA) demonstrated that DeepSignal is able to aggregate mutations on both kinase/SH2 domains and substrates to quantify binding specificity changes, predict cancer genes related to signaling transduction, and identify novel perturbed pathways.\n\nAvailabilityImplementation of DeepSignal is at https://github.com/luoyunan/DeepSignal

bioinformatics

A learning-based framework for miRNA-disease association prediction using neural networks

MotivationA microRNA (miRNA) is a type of non-coding RNA, which plays important roles in many bio-logical processes. Lots of studies have shown that miRNAs were implicated in human diseases, indicating that miRNAs might be potential biomarkers for various types of diseases. Therefore, it is important to reveal the relationships between miRNAs and diseases/phenotypes.\n\nResultsWe proposed a novel learning-based framework, MDA-CNN, for miRNA-disease association identification. The model first captures richer interaction features between diseases and miRNAs based on a three-layer network with an additional gene layer. An auto-encoder is then employed to identify the essential feature combination for each pair of miRNA and disease automatically. A final label is given by a convolutional neural network taking the reduced feature representation as input. The evaluation results showed that the proposed framework outperformed some state-of-the-art approaches in a large margin on both tasks of miRNA-disease associations prediction and miRNA-phenotype associations prediction.\n\nAvailabilityThe software will be available after the manuscript is accepted.\n\nContactjiajiepeng@nwpu.edu.cn;zywei@fudan.edu.cn;shang@nwpu.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Characterizing functional consequences of DNA copy number alterations in breast and ovarian tumors by spaceMap

MotivationWe propose a novel conditional graphical model -- spaceMap -- to construct gene regulatory networks from multiple types of high dimensional omic profiles. A motivating application is to characterize the perturbation of DNA copy number alterations (CNA) on downstream protein levels in tumors. Through a penalized multivariate regression framework, spaceMap jointly models high dimensional protein levels as responses and high dimensional CNA as predictors. In this setup, spaceMap infers an undirected network among proteins together with a directed network encoding how CNA perturb the protein network. spaceMap can be applied to learn other types of regulatory relationships from high dimensional molecular pro-files, especially those exhibiting hub structures.\n\nResultsSimulation studies show spaceMap has greater power in detecting regulatory relationships over competing methods. Additionally, spaceMap includes a network analysis toolkit for biological interpretation of inferred networks. We applied spaceMap to the CNA, gene expression and proteomics data sets from CPTAC-TCGA breast (n=77) and ovarian (n=174) cancer studies. Each cancer exhibited disruption of ion transmembrane transport and regulation from RNA polymerase II promoter by CNA events unique to each cancer. Moreover, using protein levels as a response yields a more functionally-enriched network than using RNA expressions in both cancer types. The network results also help to pinpoint crucial cancer genes and provide insights on the functional consequences of important CNA in breast and ovarian cancers.\n\nAvailabilityThe R package spaceMap -- including vignettes and documentation -- is hosted at https://topherconley.github.io/spacemap

bioinformatics

Easy Hi-C: A simple efficient protocol for 3D genome mapping in small cell populations

Despite the growing interest in studying the mammalian genome organization, it is still challenging to map the DNA contacts genome-wide. Here we present easy Hi-C (eHi-C), a highly efficient method for unbiased mapping of 3D genome architecture. The eHi-C protocol only involves a series of enzymatic reactions and maximizes the recovery of DNA products from proximity ligation. We show that eHi-C can be performed with 0.1 million cells and yields high quality libraries comparable to Hi-C.

genomics

Drosophila Fezf coordinates laminar-specific connectivity through cell-intrinsic and cell-extrinsic mechanisms

Laminar arrangement of neural connections is a fundamental feature of neural circuit organization. Identifying mechanisms that coordinate neural connections within correct layers is thus vital for understanding how neural circuits are assembled. In the medulla of the Drosophila visual system neurons form connections within ten parallel layers. The M3 layer receives input from two neuron types that sequentially innervate M3 during development. Here we show that M3-specific innervation by both neurons is coordinated by Drosophila Fezf (dFezf), a conserved transcription factor that is selectively expressed by the earlier targeting input neuron. In this cell, dFezf instructs layer specificity and activates the expression of a secreted molecule (Netrin) that regulates the layer specificity of the other input neuron. We propose that employment of transcriptional modules that cell-intrinsically target neurons to specific layers, and cell-extrinsically recruit other neurons is a general mechanism for building layered networks of neural connections.

neuroscience

Evolution of mechanisms that control mating in Drosophila males

Genetically wired neural mechanisms inhibit mating between species because even naive animals rarely mate with other species. These mechanisms can evolve through changes in expression or function of key genes in specific sensory pathways or central circuits. Gr32a is a gustatory chemoreceptor that, in D. melanogaster, is essential to inhibit interspecies courtship and sense quinine. Similar to D. melanogaster, D. simulans Gr32a is expressed in foreleg tarsi, sensorimotor appendages that inhibit interspecies courtship in both species, and it is required to sense quinine. Nevertheless, Gr32a is not required to inhibit interspecies mating by D. simulans males. However, and similar to its function in D. melanogaster, Ppk25, a member of the Pickpocket family, promotes conspecific courtship in D. simulans. Taken together, we have identified shared as well as distinct evolutionary solutions to chemosensory processing of tastants as well as cues that inhibit or promote courtship in two closely related Drosophila species.

neuroscience

Identification of Pathways Associated with Chemosensitivity through Network Embedding

Basal gene expression levels have been shown to be predictive of cellular response to cytotoxic treatments. However, such analyses do not fully reveal complex genotype-phenotype relationships, which are partly encoded in highly interconnected molecular networks. Biological pathways provide a complementary way of understanding drug response variation among individuals. In this study, we integrate chemosensitivity data from a recent pharmacogenomics study with basal gene expression data from the CCLE project and prior knowledge of molecular networks to identify specific pathways mediating chemical response. We first develop a computational method called PACER, which ranks pathways for enrichment in a given set of genes using a novel network embedding method. It examines known relationships among genes as encoded in a molecular network along with gene memberships of all pathways to determine a vector representation of each gene and pathway in the same low-dimensional vector space. The relevance of a pathway to the given gene set is then captured by the similarity between the pathway vector and gene vectors. To apply this approach to chemosensitivity data, we identify genes with basal expression levels in a panel of cell lines that are correlated with cytotoxic response to a compound, and then rank pathways for relevance to these response-correlated genes using PACER. Extensive evaluation of this approach on benchmarks constructed from databases of compound target genes, compound chemical structure, as well as large collections of drug response signatures demonstrates its advantages in identifying compound-pathway associations, compared to existing statistical methods of pathway enrichment analysis. The associations identified by PACER can serve as testable hypotheses about chemosensitivity pathways and help further study the mechanism of action of specific cytotoxic drugs. More broadly, PACER represents a novel technique of identifying enriched properties of any gene set of interest while also taking into account networks of known gene-gene relationships and interactions.

pharmacology and toxicology

GEM: A manifold learning based framework for reconstructing spatial organizations of chromosomes

Decoding the spatial organizations of chromosomes has crucial implications for studying eukaryotic gene regulation. Recently, Chromosomal conformation capture based technologies, such as Hi-C, have been widely used to uncover the interaction frequencies of genomic loci in high-throughput and genome-wide manner and provide new insights into the folding of three-dimensional (3D) genome structure. In this paper, we develop a novel manifold learning framework, called GEM (Genomic organization reconstructor based on conformational Energy and Manifold learning), to elucidate the underlying 3D spatial organizations of chromosomes from Hi-C data. Unlike previous chromatin structure reconstruction methods, which explicitly assume specific relationships between Hi-C interaction frequencies and spatial distances between distal genomic loci, GEM is able to reconstruct an ensemble of chromatin conformations by directly embedding the neigh-boring affinities from Hi-C space into 3D Euclidean space based on a manifold learning strategy that considers both the fitness of Hi-C data and the biophysical feasibility of the modeled structures, which are measured by the conformational energy derived from our current biophysical knowledge about the 3D polymer model. Extensive validation tests on both simulated interaction frequency data and experimental Hi-C data of yeast and human demonstrated that GEM not only greatly outperformed other state-of-art modeling methods but also reconstructed accurate chromatin structures that agreed well with the hold-out or independent Hi-C data and sparse geometric restraints derived from the previous fluorescence in situ hybridization (FISH) studies. In addition, as GEM can generate accurate spatial organizations of chromosomes by integrating both experimentally-derived spatial contacts and conformational energy, we for the first time extended our modeling method to recover long-range genomic interactions that are missing from the original Hi-C data. All these results indicated that GEM can provide a physically and physiologically valid 3D representations of the organizations of chromosomes and thus serve as an effective and useful genome structure reconstructor.

bioinformatics

Learning Structural Motif Representations For Efficient Protein Structure Search

MotivationUnderstanding the relationship between protein structure and function is a fundamental problem in protein science. Given a protein of unknown function, fast identification of similar protein structures from the Protein Data Bank (PDB) is a critical step for inferring its biological function. Such structural neighbors can provide evolutionary insights into protein conformation, interfaces and binding sites that are not detectable from sequence similarity. However, the computational cost of performing pairwise structural alignment against all structures in PDB is prohibitively expensive. Alignment-free approaches have been introduced to enable fast but coarse comparisons by representing each protein as a vector of structure features or fingerprints and only computing similarity between vectors. As a notable example, FragBag represents each protein by a \"bag of fragments\", which is a vector of frequencies of contiguous short backbone fragments from a predetermined library.\n\nResultsHere we present a new approach to learning effective structural motif presentations using deep learning. We develop DeepFold, a deep convolutional neural network model to extract structural motif features of a protein structure. Similar to FragBag, DeepFold represents each protein structure or fold using a vector of learned structural motif features. We demonstrate that DeepFold substantially outperforms FragBag on protein structural search on a non-redundant protein structure database and a set of newly released structures. Remarkably, DeepFold not only extracts meaningful backbone segments but also finds important long-range interacting motifs for structural comparison. We expect that DeepFold will provide new insights into the evolution and hierarchical organization of protein structural motifs.\n\nAvailabilityhttps://github.com/largelymfs/DeepFold\n\nContactjianpeng@illinois.edu

bioinformatics