Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

lncDIFF: a novel distribution-free method for differential expression analysis of long non-coding RNA

MotivationLong non-coding RNA expression data has been increasingly used in finding diagnostic and prognostic biomarkers in cancer studies. Existing differential analysis tools for RNA sequencing does not effectively accommodate low abundant genes, as commonly observed in lncRNA. We propose a novel and robust statistical method lncDIFF to detect differential expressed (DE) genes without assuming the true density on normalized counts.\n\nResultslncDIFF adopts the generalized linear model with zero-inflated exponential quasi likelihood to estimate group effect on normalized counts, and employs the likelihood ratio test to detect differential expressed genes. The proposed method and tool is suitable for data processed with standard RNA-Seq preprocessing and normalization pipelines. Simulation results illustrate that lncDIFF detects DE genes with more power and lower false discovery rate regardless of the data pattern. The analysis on a head and neck squamous cell carcinomas study also confirms that lncDIFF has better sensitivity in identifying novel lncRNA genes with relatively large fold change and prognostic value.\n\nAvailability and ImplementationlncDIFF is an R package available at https://github.com/qianli10000/lncDIFF.\n\nSupplementary InformationSupplementary Data are available at Bioinformatics online.

bioinformatics

A new statistic for efficient detection of repetitive sequences

Detecting sequences containing repetitive regions is a basic bioinformatics task with many applications. Several methods have been developed for various types of repeat detection tasks. An efficient generic method for detecting all types of repetitive sequences is still desirable.\n\nInspired by the excellent properties and successful applications of the D2 family of statistics in comparative analyses of genomic sequences, we developed a new statistic [Formula] that can efficiently discriminate sequences with or without repetitive regions. Using the statistic, we developed an algorithm of linear complexity in both computation time and memory usage for detecting all types of repetitive sequences in multiple scenarios, including finding candidate CRISPR regions from bacterial genomic or metagenomics sequences. Simulation and real data experiments showed that the method works well on both assembled sequences and unassembled short reads.

bioinformatics

rMETL: sensitive and fast mobile element insertion detection with long read realignment

SummaryMobile element insertion (MEI) is a major category of structure variations (SVs). The rapid development of long read sequencing provides the opportunity to sensitively discover MEIs. However, the signals of MEIs implied by noisy long reads are highly complex, due to the repetitiveness of mobile elements as well as the serious sequencing errors. Herein, we propose Realignment-based Mobile Element insertion detection Tool for Long read (rMETL). rMETL takes advantage of its novel chimeric read re-alignment approach to well handle complex MEI signals. Benchmarking results on simulated and real datasets demonstrated that rMETL has the ability to more sensitivity discover MEIs as well as prevent false positives. It is suited to produce high quality MEI callsets in many genomics studies.\n\nAvailability and Implementation: rMETL is available from https://github.com/hitbc/rMETL.\n\nContact: ydwang@hit.edu.cn\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

CANCERSIGN: a user-friendly and robust tool for identification and classification of mutational signatures and patterns in cancer genomes

Analyses of large somatic mutation datasets, using advanced computational algorithms, have revealed at least 30 independent mutational signatures in tumor samples. These studies have been instrumental in identification and quantification of responsible endogenous and exogenous molecular processes against cancer. The quantitative approach used to deconvolute mutational signatures is becoming an integral part of cancer research. Therefore, development of a stand-alone tool with a user-friendly graphical interface for analysis of cancer mutational signatures is necessary. In this manuscript, we introduce CANCERSIGN as an open access1 bioinformatics tool that uses raw mutation data (BED files) as input, and generates 3-mer and 5-mer mutational signatures. Additionally, this tool enables users to perform clustering on tumor samples based on the raw mutation counts as well as using the proportion of mutational signatures in each sample. Using this tool, we analysed all the whole genome somatic mutation datasets of International Cancer Genome Consortium (ICGC) samples and identified a number of novel signatures.

bioinformatics

MetaMap, an interactive webtool for the exploration of metatranscriptomic reads in human disease-related RNA-seq data

MotivationThe MetaMap resource contains metatranscriptomic expression data from screening >17,000 RNA-seq samples from >400 archived human disease-related studies for viral and microbial reads, so-called \"metafeatures\". However, navigating this set of large and heterogeneous data is challenging, especially for researchers without bioinformatic expertise. Therefore, a user-friendly interface is needed that allows users to visualize and statistically analyse the data.\n\nResultsWe developed an interactive frontend to facilitate the exploration of the MetaMap resource. The webtool allows users to query the resource by searching study abstracts for keywords or browsing expression patterns for specific metafeatures. Moreover, users can manually define sample groupings or use the existing annotation for downstream analysis. The web tool provides a large variety of analyses and visualizations including dimension reduction, differential abundance analysis and Krona visualizations. The MetaMap webtool represents a valuable resource for hypothesis generation regarding the impact of the microbiome in human disease.\n\nAvailabilityThe presented web tool can be accessed at https://github.com/theislab/MetaMap

bioinformatics

Protein model quality assessment using 3D oriented convolutional neural networks

Protein model quality assessment (QA) is a crucial and yet open problem in structural bioinformatics. The current best methods for single-model QA typically combine results from different approaches, each based on different input features constructed by experts in the field. Then, the prediction model is trained using a machine-learning algorithm. Recently, with the development of convolutional neural networks (CNN), the training paradigm has changed. In computer vision, the expert-developed features have been significantly overpassed by automatically trained convolutional filters. This motivated us to apply a three-dimensional (3D) CNN to the problem of protein model QA.\n\nWe developed a novel method for single-model QA called Ornate. Ornate (Oriented Routed Neural network with Automatic Typing) is a residue-wise scoring function that takes as input 3D density maps. It predicts the local (residue-wise) and the global model quality through a deep 3D CNN. Specifically, Ornate aligns the input density map, corresponding to each residue and its neighborhood, with the backbone topology of this residue. This circumvents the problem of ambiguous orientations of the initial models. Also, Ornate includes automatic identification of atom types and dynamic routing of the data in the network. Established benchmarks (CASP 11 and CASP 12) demonstrate the state-of-the-art performance of our approach among singlemodel QA methods.\n\nThe method is available at https://team.inria.fr/nanod/software/Ornate/. It consists of a C++ executable that transforms molecular structures into volumetric density maps, and a Python code based on the TensorFlow framework for applying the Ornate model to these maps.

bioinformatics

NetTCR: sequence-based prediction of TCR binding to peptide-MHC complexes using convolutional neural networks

Predicting epitopes recognized by cytotoxic T cells has been a long standing challenge within the field of immuno- and bioinformatics. While reliable predictions of peptide binding are available for most Major Histocompatibility Complex class I (MHCI) alleles, prediction models of T cell receptor (TCR) interactions with MHC class I-peptide complexes remain poor due to the limited amount of available training data. Recent next generation sequencing projects have however generated a considerable amount of data relating TCR sequences with their cognate HLA-peptide complex target. Here, we utilize such data to train a sequence-based predictor of the interaction between TCRs and peptides presented by the most common human MHCI allele, HLA-A*02:01. Our model is based on convolutional neural networks, which are especially designed to meet the challenges posed by the large length variations of TCRs. We show that such a sequence-based model allows for the identification of TCRs binding a given cognate peptide-MHC target out of a large pool of non-binding TCRs.

bioinformatics

Design of an epitope-based peptide vaccine against Cryptococcus neoformans

IntroductionThis study aimed to design an immunogenic epitope for Cryptococcus neoformans the etiological agent of cryptococcosis using in silico simulations, for epitope prediction, we selected the mannoprotein antigen MP88 which its known to induce protective immunity.\n\nMaterial & methodA total of 39 sequences of MP88 protein with length 378 amino acids were retrieved from the National Center for Biotechnology Information database (NCBI) in the FASTA format were used to predict antigenic B-cell and T cell epitopes via different bioinformatics tools at Immune Epitope Database and Analysis Resource (IEDB). The tertiary structure prediction of MP88 was created in RaptorX, and visualized by UCSF Chimera software.\n\nResultA Conserved B-cell epitopes AYSTPA, AYSTPAS, PASSNCK, and DSAYPP have displayed the most promising B cell epitopes. While the YMAADQFCL, VSYEEWMNY and FQQRYTGTF they represent the best candidates T-cell conserved epitopes, the 9-mer epitope YMAADQFCL display the greater interact with 9 MHC-I alleles and HLA-A*02:01 alleles have the best interaction with an epitope. The VSYEEWMNY and FQQRYTGTF they are non-allergen while YMAADQFCL was an allergen. For MHC class II peptide binding prediction, the YARLLSLNA, ISYGTAMAV and INQTSYARL represent the most Three highly binding affinity core epitopes. The core epitope INQTSYARL was found to interact with 14 MHC-II. The allergenicity prediction reveals ISYGTAMAV, INQTSYARL were non-allergen and YARLLSLNA was an allergen. Regarding population coverage the YMAADQFCL exhibit, a higher percentage among the world (69.75%) and the average population coverage was 93.01%. In MHC-II, ISYGTAMAV epitope reveal a higher percentage (74.39%) and the average population coverage was (81.94%). This successfully designed a peptide vaccine against Cryptococcus neoformans open up a new horizon in Cryptococcus neoformans research; the results require validation by in vitro and in vivo experiments.

bioinformatics

Improving on hash-based probabilistic sequence classification using multiple spaced seeds and multi-index Bloom filters

Alignment-free classification of sequences against collections of sequences has enabled high-throughput processing of sequencing data in many bioinformatics analysis pipelines. Originally hash-table based, much work has been done to improve and reduce the memory requirement of indexing of k-mer sequences with probabilistic indexing strategies. These efforts have led to lower memory highly efficient indexes, but often lack sensitivity in the face of sequencing errors or polymorphism because they are k-mer based. To address this, we designed a new memory efficient data structure that can tolerate mismatches using multiple spaced seeds, called a multi-index Bloom Filter. Implemented as part of BioBloom Tools, we demonstrate our algorithm in two applications, read binning for targeted assembly and taxonomic read assignment. Our tool shows a higher sensitivity and specificity for read-binning than BWA MEM at an order of magnitude less time. For taxonomic classification, we show higher sensitivity than CLARK-S at an order of magnitude less time while using half the memory.

bioinformatics

EMBL2checklists: A Python package to facilitate the user-friendly submission of plant DNA barcoding sequences to ENA

BackgroundThe submission of DNA sequences to public sequence databases is an essential, but insufficiently automated step in the process of generating and disseminating novel DNA sequence data. Despite the centrality of database submissions to biological research, the range of available software tools that facilitate the preparation of sequence data for database submissions is low, especially for sequences generated via plant DNA barcoding. Current submission procedures can be complex and prohibitively time expensive for any but a small number of input sequences. A user-friendly software tool is needed that streamlines the file preparation for database submissions of DNA sequences that are commonly generated in plant DNA barcoding.\n\nMethodsA Python package was developed that converts DNA sequences from the common EMBL and GenBank flat file formats to submission-ready, tab-delimited spreadsheets (so-called \"checklists\") for a subsequent upload to the public sequence database of the European Nucleotide Archive (ENA). The software tool, titled \"EMBL2checklists\", automatically converts DNA sequences, their annotation features, and associated metadata into the idiosyncratic format of marker-specific ENA checklists and, thus, generates output that can be uploaded via the interactive Webin submission system of ENA.\n\nResultsEMBL2checklists provides a simple, platform-independent tool that automates the conversion of common plant DNA barcoding sequences into easily editable spreadsheets that require no further processing but their upload to ENA via the interactive Webin submission system. The software is equipped with an intuitive graphical as well as an efficient command-line interface for its operation. The utility of the software is illustrated by its application in the submission of DNA sequences of two recent plant phylogenetic investigations and one fungal metagenomic study.\n\nDiscussionEMBL2checklists bridges the gap between common software suites for DNA sequence assembly and annotation and the interactive data submission process of ENA. It represents an easy-to-use solution for plant biologists without bioinformatics expertise to generate submission-ready checklists from common plant DNA sequence data. It allows the post-processing of checklists as well as work-sharing during the submission process and solves a critical bottleneck in the effort to increase participation in public data sharing.

bioinformatics

Comprehensive in silico Analysis of IKBKAP gene that could potentially cause Familial dysautonomia

BackgroundFamilial dysautonomia (FD) is a rare neurodevelopmental genetic disorder within the larger classification of hereditary sensory and autonomic neuropathies. We aimed to identify the pathogenic SNPs in IKBKAP gene by computational analysis softwares, and to determine the structure, function and regulation of their respective proteins.\n\nMaterials and MethodsWe carried out in silico analysis of structural effect of each SNP using different bioinformatics tools to predict SNPs influence on protein structure and function.\n\nResult41 novel mutations out of 973 nsSNPs that are found be deleterious effect on the IKBKAP structure and function.\n\nConclusionThis is the first in silico analysis in IKBKAP gene to prioritize SNPs for further genetic studies.

bioinformatics

AIDE: annotation-assisted isoform discovery and abundanceestimation from RNA-seq data

Genome-wide accurate identification and quantification of full-length mRNA isoforms is crucial for investigating transcriptional and post-transcriptional regulatory mechanisms of biological phenomena. Despite continuing efforts in developing effective computational tools to identify or assemble full-length mRNA isoforms from second-generation RNA-seq data, it remains a challenge to accurately identify mRNA isoforms from short sequence reads due to the substantial information loss in RNA-seq experiments. Here we introduce a novel statistical method, AIDE (Annotation-assisted Isoform DiscovEry), the first approach that directly controls false isoform discoveries by implementing the testing-based model selection principle. Solving the isoform discovery problem in a stepwise and conservative manner, AIDE prioritizes the annotated isoforms and precisely identifies novel isoforms whose addition significantly improves the explanation of observed RNA-seq reads. We evaluate the performance of AIDE based on multiple simulated and real RNA-seq datasets followed by a PCR-Sanger sequencing validation. Our results show that AIDE effectively leverages the annotation information to compensate the information loss due to short read lengths. AIDE achieves the highest precision in isoform discovery and the lowest error rates in isoform abundance estimation, compared with three state-of-the-art methods Cufflinks, SLIDE, and StringTie. As a robust bioinformatics tool for transcriptome analysis, AIDE will enable researchers to discover novel transcripts with high confidence.

bioinformatics

Using the drug-protein interactome to identify anti-ageing compounds for humans

Advancing age is the dominant risk factor for most of the major killer diseases in developed countries. Hence, ameliorating the effects of ageing may prevent multiple diseases simultaneously. Drugs licensed for human use against specific diseases have proved to be effective in extending lifespan and healthspan in animal models, suggesting that there is scope for drug repurposing in humans. New bioinformatic methods to identify and prioritise potential anti-ageing compounds for humans are therefore of interest. In this study, we first used drug-protein interaction information, to rank 1,147 drugs by their likelihood of targeting ageing-related gene products in humans. Among 19 statistically significant drugs, 6 have already been shown to have pro-longevity properties in animal models (p < 0.001). Using the targets of each drug, we established its association with ageing at multiple levels of biological actions including pathways, functions and protein interactions. Finally, combining all the data, we calculated a comprehensive ranked list of drugs that predicted tanespimycin, an inhibitor of HSP-90, as the top-ranked novel anti-ageing candidate. We experimentally validated the pro-longevity effect of tanespimycin through its HSP-90 target in Caenorhabditis elegans.\n\nAuthor SummaryHuman life expectancy is continuing to increase worldwide, as a result of successive improvements in living conditions and medical care. Although this trend is to be celebrated, advancing age is the major risk factor for multiple impairments and chronic diseases. As a result, the later years of life are often spent in poor health and lowered quality of life. However, these effects of ageing are not inevitable, because very long-lived people often suffer rather little ill-health at the end of their lives. Furthermore, laboratory experiments have shown that animals fed with specific drugs can live longer and with fewer age-related diseases than their untreated companions. We therefore need to identify drugs with anti-ageing properties for humans. We have therefore used computers to search for drugs that affect components and processes known to be important in human ageing. This approach worked, because it was able to re-discover several drugs known to increase lifespan in animal models, plus some new ones, including one that we tested experimentally and validated in this study. These drugs are now a high priority for animal testing and for exploring effects on human ageing.

bioinformatics

Quantifying point-mutations in shotgun metagenomic data

Metagenomics has emerged as a central technique for studying the structure and function of microbial communities. Often the functional analysis is restricted to classification into broad functional categories. However, important phenotypic differences, such as resistance to antibiotics, are often the result of just one or a few point mutations in otherwise identical sequences. Bioinformatic methods for metagenomic analysis have generally been poor at accounting for this fact, resulting in a somewhat limited picture of important aspects of microbial communities. Here, we address this problem by providing a software tool called Mumame, which can distinguish between wildtype and mutated sequences in shotgun metagenomic data and quantify their relative abundances. We demonstrate the utility of the tool by quantifying antibiotic resistance mutations in several publicly available metagenomic data sets. We also identified that sequencing depth is a key factor to detect rare mutations. Therefore, much larger numbers of sequences may be required for reliable detection of mutations than for most other applications of shotgun metagenomics. Mumame is freely available from http://microbiology.se/software/mumame

bioinformatics

AMON: Annotation of metabolite origins via networks to better integrate microbiome and metabolome data

MotivationUntargeted metabolomics of host-associated samples has yielded insights into mechanisms by which microbes modulate health. However, data interpretation is challenged by the complexity of origins of the small molecules measured, which can come from the host, microbes that live with the host, or from other exposures such as diet or the environment.\n\nResultsWe address this challenge through development of AMON: Annotation of Metabolite Origins via Networks. AMON is an open-source bioinformatics application that can be used to determine the degree to which annotated compounds in the metabolome may have been produced by bacteria present, the host, either (i.e. both the bacteria and host are capable of production), or neither (i.e. neither the human or the fecal microbiome are predicted to be capable of producing the observed metabolite).\n\nAvailability and ImplementationThis software is available at https://github.com/lozuponelab/AMON as well as via pip.\n\nContactcatherine.lozupone@ucdenver.edu

bioinformatics

Tibanna: software for scalable execution of portable pipelines on the cloud

SummaryWe introduce Tibanna, an open-source software tool for automated execution of bioinformatics pipelines on Amazon Web Services (AWS). Tibanna accepts reproducible and portable pipeline standards including Common Workflow Language (CWL), Workflow Description Language (WDL) and Docker. It adopts a strategy of isolation and optimization of individual executions, combined with a serverless scheduling approach. Pipelines are executed and monitored using local commands or the Python Application Programming Interface (API) and cloud configuration is automatically handled. Tibanna is well suited for projects with a range of computational requirements, including those with large and widely fluctuating loads. Notably, it has been used to process terabytes of data for the 4D Nucleome (4DN) Network.\n\nAvailabilitySource code is available on GitHub at https://github.com/4dn-dcic/tibanna.

bioinformatics

A framework for space-efficient variable-order Markov models

MotivationMarkov models with contexts of variable length are widely used in bioinformatics for representing sets of sequences with similar biological properties. When models contain many long contexts, existing implementations are either unable to handle genome-scale training datasets within typical memory budgets, or they are optimized for specific model variants and are thus inflexible.\n\nResultsWe provide practical, versatile representations of variable-order Markov models and of interpolated Markov models, that support a large number of context-selection criteria, scoring functions, probability smoothing methods, and interpolations, and that take up to 4 times less space than previous implementations based on the suffix array, regardless of the number and length of contexts, and up to 10 times less space than previous trie-based representations, or more, while matching the size of related, state-of-the-art data structures from Natural Language Processing. We describe how to further compress our indexes to a quantity related to the redundancy of the training data, saving up to 90% of their space on repetitive datasets, and making them become up to 60 times smaller than previous implementations based on the suffix array. Finally, we show how to exploit constraints on the length and frequency of contexts to further shrink our compressed indexes to half of their size or more, achieving data structures that are 100 times smaller than previous implementations based on the suffix array, or more. This allows variable-order Markov models to be trained on bigger datasets and with longer contexts on the same hardware, thus possibly enabling new applications.\n\nAvailability and implementationhttps://github.com/jnalanko/VOMM

bioinformatics

netome: a computational framework for metabolite profiling and omics network analysis

SummaryAdvances in metabolomics technologies have enabled comprehensive analyses of associations between metabolites and human disease and have provided a means to study biochemical pathways and processes in detail using model systems. Liquid chromatography tandem mass spectrometry (LC-MS) is an analytical technique commonly used by metabolomics labs to measure hundreds of metabolites of known identity and thousands of \"peaks\" from yet to be identified compounds that are tracked by their measured masses and chromatographic retention times. netome is a computational framework that provides tools for analyzing processed LC-MS data. In this framework, we develop and provide various computational resources including individual software modules to inspect and adjust trends in raw data, align unknown peaks between separately acquired data sets, and to remove redundancies in nontargeted LC-MS data arising from multiple ionization products of a single metabolite. These tools are deployed through computing resources such as web servers and virtual machines with detailed documentation in order to support researchers.\n\nAvailability and implementationnetome is publicly available with extensive documentation and support via issue tracker at https://broadinstitute.github.io/netome under the MIT license. netome includes a set of computational methods that have been designed to execute quality control and post-raw data processing tasks for metabolomics data (e.g. scaling and clustering metabolite abundances), as well as statistical association testing in a network manner (e.g., testing relationship between metabolites and microbes). Each individual tool is available with source code, workshop-oriented documentation which includes instructions for installation and using tools with demonstration examples, and a web server with all services. We also provide a complete image of the netome package with all pre-installed dependencies and support for Google Compute Engine and Amazon EC2. All tools and related services are maintained, and upon new developments, new modules will be added to the environment.\n\nContactrah@broadinstitute.org, clary@broadinstitute.org\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics