Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Fast Nonnegative Matrix Factorization andApplications to Pattern Extraction, Deconvolutionand Imputation

Nonnegative matrix factorization (NMF) is a technique widely used in various fields, including artificial intelligence (AI), signal processing and bioinformatics. However existing algorithms and R packages cannot be applied to large matrices due to their slow convergence, and cannot handle missing values. In addition, most NMF research focuses only on blind decompositions: decomposition without utilizing prior knowledge. We adapt the idea of sequential coordinate-wise descent to NMF to increase the convergence rate. Our NMF algorithm thus handles missing values naturally and integrates prior knowledge to guide NMF towards a more meaningful decomposition. To support its use, we describe a novel imputation-based method to determine the rank of decomposition. All our algorithms are implemented in the R package NNLM, which is freely available on CRAN.

bioinformatics

Inferring Edge Function in Protein-Protein Interaction Networks

Motivation: Post-translational modifications (PTMs) regulate many key cellular processes. Numerous studies have linked the topology of protein-protein interaction (PPI) networks to many biological phenomena such as key regulatory processes and disease. However, these methods fail to give insight in the functional nature of these interactions. On the other hand, pathways are commonly used to gain biological insight into the function of PPIs in the context of cascading interactions, sacrificing the coverage of networks for rich functional annotations on each PPI. We present a machine learning approach that uses Gene Ontology, InterPro and Pfam annotations to infer the edge functions in PPI networks, allowing us to combine the high coverage of networks with the information richness of pathways.\n\nResults: An ensemble method with a combination Logistic Regression and Random Forest classifiers trained on a high-quality set of annotated interactions, with a total of 18 unique labels, achieves high a average F1 score 0.88 despite not taking advantage of multi-label dependencies. When applied to the human interactome, our method confidently classifies 62% of interactions at a probability of 0.7 or higher.\n\nAvailability: Software and data are available at https://github.com/DavisLaboratory/pyPPI\n\nContact: davis.m@wehi.edu.au\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

An average-case sublinear exact Li and Stephens forward algorithm

Hidden Markov models of haplotype inheritance such as the Li and Stephens model allow for computationally tractable probability calculations using the forward algorithms as long as the representative reference panel used in the model is sufficiently small. Specifically, the monoploid Li and Stephens model and its variants are linear in reference panel size unless heuristic approximations are used. However, sequencing projects numbering in the thousands to hundreds of thousands of individuals are underway, and others numbering in the millions are anticipated.\n\nTo make the Li and Stephens forward algorithm for these datasets computationally tractable, we have created a numerically exact version of the algorithm with observed average case [O](nk0.35) runtime, avoiding any tradeoff between runtime and model complexity. We demonstrate that our approach also provides a succinct data structure for general purpose haplotype data storage. We discuss generalizations of our algorithmic techniques to other hidden Markov models.\n\n2012 ACM Subject ClassificationTheory of computation {longrightarrow} Streaming, sublinear and near linear time algorithms; Applied computing {longrightarrow} Bioinformatics\n\nSupplement Materialhttps://github.com/yoheirosen/sublinear-Li-Stephens.\n\nFundingThis work was supported by the National Human Genome Research Institute of the National Institutes of Health under Award Number 5U54HG007990, the National Heart, Lung, and Blood Institute of the National Institutes of Health under Award Number 1U01HL137183-01, and grants from the W.M. Keck foundation and the Simons Foundation.\n\nAcknowledgementsWe would like to thank Jordan Eizenga for his helpful discussions throughout the development of this work.

bioinformatics

GRIMM: GRaph IMputation and Matching for HLA Genotypes

Motivation: For over 10 years allele-level HLA matching for bone marrow registries has been performed in a probabilistic context. HLA typing technologies provide ambiguous results in that they could not distinguish among all known HLA allele sequences, therefore registries have implemented matching algorithms that provide lists of donor and cord blood units ordered in terms of the likelihood of allele-level matching at specific HLA loci. With the growth of registry sizes, current match algorithm implementations are unable to provide match results in real time.\n\nResults: We present here novel computationally-efficient open source implementation of an HLA imputation and match algorithm using a graph database platform. Using graph traversal, our algorithm runtime grows slowly with registry size. This implementation generates results that agree with consensus output on a publicly-available match algorithm crossvalidation dataset.\n\nAvailability: The Python, Perl and Neo4jJcode is available at https://git.com/nmdp-bioinformatics/grimm\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

Construction and analysis of mRNA, miRNA, lncRNA, and TF regulatory networks reveal the key genes in prostate cancer

Purpose: Prostate cancer (PCa) causes a common male urinary system malignant tumour, and the molecular mechanisms of PCa remain poorly understood. This study aims to investigate the underlying molecular mechanisms of PCa with bioinformatics.\n\nMethods: Original gene expression profiles were obtained from the GSE64318 and GSE46602 datasets in the Gene Expression Omnibus (GEO). We conducted differential screens of the expression of genes (DEGs) between two groups using the R software limma package. The interactions between the differentially expressed miRNAs, mRNAs and lncRNAs were predicted and merged with the target genes. Co-expression of the miRNAs, lncRNAs and mRNAs were selected to construct the mRNA-miRNA and-lncRNA interaction networks. Gene Ontology (GO) and Kyoto Encyclopaedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed for the DEGs. The protein-protein interaction (PPI) networks were constructed, and the transcription factors were annotated. The expression of hub genes in the TCGA datasets was verified to improve the reliability of our analysis.\n\nResults: The results demonstrated that 60 miRNAs, 1578 mRNAs and 61 lncRNAs were differentially expressed in PCa. The mRNA-miRNA-lncRNA networks were composed of 5 miRNA nodes, 13 lncRNA nodes, and 45 mRNA nodes. The DEGs were mainly enriched in the nuclei and cytoplasm and were involved in the regulation of transcription, related to sequence-specific DNA binding, and participated in the regulation of the PI3K-Akt signalling pathway. These pathways are related to cancer and focal adhesion signalling pathways. Furthermore, we found that 5 miRNAs, 6 lncRNAs, 6 mRNAs and 2 TFs play important regulatory roles in the interaction network. The expression levels of EGFR, VEGFA, PIK3R1, DLG4, TGFBR1 and KIT were significantly different between PCa and normal prostate tissue.\n\nConclusion: Based on the current study, large-scale effects of interrelated mRNAs, miRNAs, lncRNAs, and TFs were revealed and a model for predicting the mechanism of PCa was provided. This study provides new insight for the exploration of the molecular mechanisms of PCa and valuable clues for further research.

bioinformatics

Re-assembly, quality evaluation, and annotation of 678 microbial eukaryotic reference transcriptomes

BackgroundDe novo transcriptome assemblies are required prior to analyzing RNAseq data from a species without an existing reference genome or transcriptome. Despite the prevalence of transcriptomic studies, the effects of using different workflows, or \"pipelines\", on the resulting assemblies are poorly understood. Here, a pipeline was programmatically automated and used to assemble and annotate raw transcriptomic short read data collected by the Marine Microbial Eukaryotic Transcriptome Sequencing Project (MMETSP). The resulting transcriptome assemblies were evaluated and compared against assemblies that were previously generated with a different pipeline developed by the National Center for Genome Research (NCGR).\n\nResultsNew transcriptome assemblies contained the majority of previous contigs as well as new content. On average, 7.8% of the annotated contigs in the new assemblies were novel gene names not found in the previous assemblies. Taxonomic trends were observed in the assembly metrics, with assemblies from the Dinoflagellata and Ciliophora phyla showing a higher percentage of open reading frames and number of contigs than transcriptomes from other phyla.\n\nConclusionsGiven current bioinformatics approaches, there is no single best reference transcriptome for a particular set of raw data. As the optimum transcriptome is a moving target, improving (or not) with new tools and approaches, automated and programmable pipelines are invaluable for managing the computationally-intensive tasks required for re-processing large sets of samples with revised pipelines and ensuring a common evaluation workflow is applied to all samples. Thus, re-assembling existing data with new tools using automated and programmable pipelines may yield more accurate identification of taxon-specific trends across samples in addition to novel and useful products for the community.\n\nKey PointsO_LIRe-assembly with new tools can yield new results\nC_LIO_LIAutomated and programmable pipelines can be used to process arbitrarily many samples.\nC_LIO_LIAnalyzing many samples using a common pipeline identifies taxon-specific trends.\nC_LI

bioinformatics

VCPA: genomic variant calling pipeline and data management tool for Alzheimer’s Disease Sequencing Project

Summary: We report VCPA, our SNP/Indel Variant Calling Pipeline and data management tool used for analysis of whole genome and exome sequencing (WGS/WES) for the Alzheimers Disease Sequencing Project. VCPA consists of two independent but linkable components: pipeline and tracking database. The pipeline is coded in Workflow Description Language and is fully optimized for the Amazon elastic compute cloud environment. This includes steps for processing raw sequence reads including read alignment, and all the way up to variant calling using GATK. The tracking database allows users to dynamically view the statuses of jobs running and the quality metrics reported by the pipeline. Users can thus monitor the production process and diagnose if any problem arises during the procedure. All quality metrics (>100 collected per processed genome) are stored in the database, thus facilitating users to compare, share and visualize the results. To summarize, VCPA is functional equivalent to the CCDG/TOPMed pipeline. Together with the dockerized database (also available as Amazon Machine Image), users can easily process any WGS/WES data on Amazon cloud with minimal installation.\n\nAvailability: VCPA is released under the MIT license and is available for academic and nonprofit use for free. The pipeline source code and step-by-step instructions are available from the National Institute on Aging Genetics of Alzheimers Disease Data Storage Site (http://www.niagads.org/VCPA).\n\nContact: yyee@pennmedicine.upenn.edu or lswang@pennmedicine.upenn.edu\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

A Deep Learning Approach for Learning Intrinsic Protein-RNA Binding Preferences

MotivationThe complexes formed by binding of proteins to RNAs play key roles in many biological processes, such as splicing, gene expression regulation, translation, and viral replication. Understanding protein-RNA binding may thus provide important insights to the functionality and dynamics of many cellular processes. This has sparked substantial interest in exploring protein-RNA binding experimentally, and predicting it computationally. The key computational challenge is to efficiently and accurately infer RNA-binding models that will enable prediction of novel protein-RNA interactions to additional transcripts of interest.\n\nResultsWe developed DLPRB, a new deep neural network (DNN) approach for learning protein-RNA binding preferences and predicting novel interactions. We present two different network architectures: a convolutional neural network (CNN), and a recurrent neural network (RNN). The novelty of our network hinges upon two key aspects: (i) the joint analysis of both RNA sequence and structure, which is represented as a probability vector of different RNA structural contexts; (ii) novel features in the architecture of the networks, such as the application of RNNs to RNA-binding prediction, and the combination of hundreds of variable-length filters in the CNN. Our results in inferring accurate RNA-binding models from high-throughput in vitro data exhibit substantial improvements, compared to all previous approaches for protein-RNA binding prediction (both DNN and non-DNN based). A highly significant improvement is achieved for in vitro binding prediction, and a more modest, yet statistically significant,improvement for in vivo binding prediction. When incorporating experimentally-measured RNA structure compared to predicted one, the improvement on in vivo data increases. By visualizing the binding specificities, we can gain novel biological insights underlying the mechanism of protein RNA-binding.\n\nAvailabilityThe source code is publicly available at https://github.com/ilanbb/dlprb.\n\nContactyaronore@bgu.ac.il\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Homology-based loop modelling yields more complete crystallographic protein structures

Inherent protein flexibility, poor or low-resolution diffraction data, or poor electron density maps, often inhibit building complete structural models during X-ray structure determination. However, advances in crystallographic refinement and model building nowadays often allow to complete previously missing parts. Here, we present algorithms that identify regions missing in a certain model but present in homologous structures in the Protein Data Bank (PDB), and \"graft\" these regions of interest. These new regions are refined and validated in a fully automated procedure. Including these developments in our PDB-REDO pipeline, allowed to build 24,962 missing loops in the PDB. The models and the automated procedures are publically available through the PDB-REDO databank and web server (https://pdb-redo.eu). More complete protein structure models enable a higher quality public archive, but also a better understanding of protein function, better comparison between homologous structures, and more complete data mining in structural bioinformatics projects.\n\nSynopsisThousands of missing regions in existing protein structure models are completed using new methods based on homology.

bioinformatics

Solving scaffolding problem with repeats

One of the most important steps in genome assembly is scaffolding. Increasing the length of sequencing reads allows assembling short genomes but assembly of long repeat-rich genomes remains one of the most interesting and challenging problems in bioinformatics. There is a high demand in developing computational approaches for repeat aware scaffolding. In this paper, we propose a novel repeat-aware scaffolder BATISCAF based on the optimization formulation for filtering out repeated and short contigs. Our experiments with five benchmarking datasets show that the proposed tool BATISCAF outperforms state-of-the-art tools. BATISCAF is freely available on GitHub: https://github.com/mandricigor/batiscaf.

bioinformatics

ERVcaller: Identifying and genotyping non-reference unfixed endogenous retroviruses (ERVs) and other transposable elements (TEs) using next-generation sequencing data

MotivationApproximately 8% of the human genome is derived from endogenous retroviruses (ERVs). In recent years, an increasing number of human diseases have been found to be associated with ERVs. However, it remains challenging to accurately detect the full spectrum of polymorphic (unfixed) ERVs using next-generation sequencing (NGS) data.\n\nResultsWe designed a new tool, ERVcaller, to detect and genotype transposable element (TE) insertions, including ERVs, in the human genome. We evaluated ERVcaller using both simulated and real benchmark whole-genome sequencing (WGS) datasets. By comparing with existing tools, ERVcaller consistently obtained both the highest sensitivity and precision for detecting simulated ERV and other TE insertions derived from real polymorphic TE sequences. For the WGS data from the 1000 Genomes Project, ERVcaller detected the largest number of TE insertions per sample based on consensus TE loci. By analyzing the experimentally verified TE insertions, ERVcaller had 94.0% TE detection sensitivity and 96.6% genotyping accuracy. PCR and Sanger sequencing in a small sample set verified 86.7% of examined insertion statuses and 100% of examined genotypes. In conclusion, ERVcaller is capable of detecting and genotyping TE insertions using WGS data with both high sensitivity and precision. This tool can be applied broadly to other species.\n\nAvailabilitywww.uvm.edu/genomics/software/ERVcaller.html\n\nContactdawei.li@uvm.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

VARAN-GIE: Curation of Genomic Interval Sets

Genomic interval sets are fundamental elements of genome annotation and are the output of countless bioinformatics applications. Nevertheless, tool support for the manual curation of these data is currently limited. We developed VARAN-GIE, an extension of the popular Integrative Genomics Viewer (IGV) that adds functionality to edit, annotate and merge genomic interval sets. Data can easily be shared with other users and imported/exported from/to multiple common data formats. VARAN-GIE binary releases, source-code, user guides and tutorials are available at https://github.com/popitsch/varan-gie/

bioinformatics

Mapping the genetics of neuropsychological traits to the molecular network of the human brain using a data integrative approach

MotivationComplex neuropsychiatric conditions including autism spectrum disorders are among the most heritable neurodevelopmental disorders with distinct profiles of neuropsychological traits. A variety of genetic factors modulate these traits (phenotypes) underlying clinical diagnoses. To explore the associations between genetic factors and phenotypes, genome-wide association studies are broadly applied. Stringent quality checks and thorough downstream analyses for in-depth interpretation of the associations are an indispensable prerequisite. However, in the area of neuropsychology there is no framework existing, which besides performing association studies also affiliates genetic variants at the brain and gene network level within a single framework.\n\nResultsWe present a novel bioinformatics approach in the field of neuropsychology that integrates current state-of-the-art tools, algorithms and brain transcriptome data to elaborate the association of phenotype and genotype data. The integration of transcriptome data gives an advantage over the existing pipelines by directly translating genetic associations to brain regions and developmental patterns. Based on our data integrative approach, we identify genetic variants associated with Intelligence Quotient (IQ) in an autism cohort and found their respective genes to be expressed in specific brain areas.\n\nConclusionOur data integrative approach revealed that IQ is related to early down-regulated and late up-regulated gene modules implicated in frontal cortex and striatum, respectively. Besides identifying new gene associations with IQ we also provide a proof of concept, as several of the identified genes in our analysis are candidate genes related to intelligence in autism, intellectual disability, and Alzheimers disease. The framework provides a complete extensive analysis starting from a phenotypic trait data to its association at specific brain areas at vulnerable time points within a timespan of four days.\n\nAvailability and ImplementationOur framework is implemented in R and Python. It is available as an in-house script, which can be provided on demand.\n\nContactafsheen.yousaf@kgu.de

bioinformatics

Proteome-Scale Relationships Between Local Amino Acid Composition and Protein Fates and Functions

Proteins with low-complexity domains continue to emerge as key players in both normal and pathological cellular processes. Although low-complexity domains are often grouped into a single class, individual low-complexity domains can differ substantially with respect to amino acid composition. These differences may strongly influence the physical properties, cellular regulation, and molecular functions of low-complexity domains. Therefore, we developed a bioinformatic approach to explore relationships between amino acid composition, protein metabolism, and protein function. We find that local compositional enrichment within protein sequences affects the translation efficiency, abundance, half-life, subcellular localization, and molecular functions of proteins on a proteome-wide scale. However, these effects depend upon the type of amino acid enriched in a given sequence, highlighting the importance of distinguishing between different types of low-complexity domains. Furthermore, many of these effects are discernible at amino acid compositions below those required for classification as low-complexity or statistically-biased by traditional methods and in the absence of homopolymeric amino acid repeats, indicating that thresholds employed by classical methods may not reflect biologically relevant criteria. Application of our analyses to composition-driven processes, such as the formation of membraneless organelles, reveals distinct composition profiles even for closely related organelles. Collectively, these results provide a unique perspective and detailed insights into relationships between amino acid composition, protein metabolism, and protein functions.\n\nAuthor SummaryLow-complexity domains in protein sequences are regions that are composed of only a few amino acids in the protein \"alphabet\". These domains often have unique chemical properties and play important biological roles in both normal and disease-related processes.While a number of approaches have been developed to define low-complexity domains, these methods each possess conceptual limitations. Therefore, we developed a complementary approach that focuses on local amino acid composition (i.e. the amino acid composition within small regions of proteins). We find that high local composition of individual amino acids is associated with pervasive effects on protein metabolism, subcellular localization, and molecular function on a proteome-wide scale. Importantly, the nature of the effects depend on the type of amino acid enriched within the examined domains, and are observable in the absence of classically-defined low-complexity (and related) domains. Furthermore, we define the compositions of proteins involved in the formation of membraneless, protein-rich organelles such as stress granules and P-bodies. Our results provide a coherent view and unprecedented resolution of the effects of local amino acid enrichment on protein biology.

bioinformatics

Differential Expression of Immune Related Genes in Taste Buds of Fed and Fasted Mice

To study the effects of feeding on taste receptor cell transcriptional regulation, we performed RNA-seq analysis of circumvallate taste buds isolated from mice before or after food consumption. Here we report and compare the taste bud transcriptomes obtained from food-deprived, satiated, and ad libitum fed control mice. Despite sample heterogeneity inherent to the whole taste bud transcriptome, bioinformatics analysis yielded 144 differentially expressed transcripts associated with immunity, cytoskeletal structure, and protein folding between these groups. We also profiled the transcriptome obtained from ad libitum fed control mice based on receptor related gene ontology terms, demonstrating a use of this dataset in the identification of novel TRC receptors. The data presented here suggest that transcriptional regulation of immune cytokine signaling occurs in TRCs shortly after meal consumption, though additional analysis is required to confirm this hypothesis.

bioinformatics

Integrating distal and proximal information to predict gene expression via a densely con-nected convolutional neural network

MotivationInteractions among such cis-regulatory elements as enhancers and promoters are main driving forces shaping context-specific chromatin structure and gene expression. Although there have been computational methods for predicting gene expression from genomic and epigenomic information, most of them overlook long-range enhancer-promoter interactions, due to the difficulty in precisely linking regulatory enhancers to target genes. Recently, a novel high-throughput experimental approach named HiChIP has been developed and generating comprehensive data on high-resolution interactions between promoters and distal enhancers. On the other hand, plenty of studies have suggested that deep learning achieves state-of-the-art performance in epigenomic signal prediction, and thus promoting the understanding of regulatory elements. In consideration of these two factors, we integrate proximal promoter sequences and HiChIP distal enhancer-promoter interactions to accurately model gene expression.\n\nResultsWe propose DeepExpression, a densely connected convolutional neural network to predict gene expression using both promoter sequences and enhancer-promoter interactions. We demonstrate that our model consistently outperforms baseline methods not only in the classification of binary gene expression status but also in the regression of continuous gene expression levels, in both cross-validation experiments and cross-cell lines predictions. We show that sequential promoter information is more informative than experimental enhancer information while enhancer-promoter interactions are most beneficial from those within {+/-}100 kbp around the TSS of a gene. We finally visualize motifs in both promoter and enhancer regions and show the match of identified sequence signatures and known motifs. We expect to see a wide spectrum of applications using HiChIP data in deciphering the mechanism of gene regulation.\n\nAvailabilityDeepExpression is freely available at https://github.com/wanwenzeng/DeepExpression.\n\nContactruijiang@tsinghua.edu.cn, ywang@amss.ac.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

SVCollector: Optimized sample selection for validating and long-read resequencing of structural variants

SummaryStructural Variations (SVs) are increasingly recognized for their importance in genomics. Short-read sequencing is the most widely-used approach for genotyping large numbers of samples for SVs but suffers from relatively poor accuracy. Here we present SVCollector, an open-source method that optimally selects samples to maximize variant discovery and validation using long read resequencing or PCR-based validation. SVCollector has two modes: selecting those samples that are individually the most diverse or those that collectively capture the largest number of variations.\n\nAvailabilityhttps://github.com/fritzsedlazeck/SVCollector\n\nContactfritz.sedlazeck@bcm.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

A Visualization Tool to Evaluate Pairwise Protein Structure Alignment Algorithms

The alignment of two protein structures is a fundamental problem in structural bioinformatics. In this paper, we propose a novel approach to measure the effectiveness of a sample of three such algorithms, DALI, TM-align and EDAlignsse. The underlying premise of our approach is that structural proximity should translate into spatial proximity.

bioinformatics