Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

ASaiM: a Galaxy-based framework to analyze raw shotgun data from microbiota

BackgroundNew generation of sequencing platforms coupled to numerous bioinformatics tools has led to rapid technological progress in metagenomics and metatranscriptomics to investigate complex microorganism communities. Nevertheless, a combination of different bioinformatic tools remains necessary to draw conclusions out of microbiota studies. Modular and user-friendly tools would greatly improve such studies.\n\nFindingsWe therefore developed ASaiM, an Open-Source Galaxy-based framework dedicated to microbiota data analyses. ASaiM provides a curated collection of tools to explore and visualize taxonomic and functional information from raw amplicon, metagenomic or metatranscriptomic sequences. To guide different analyses, several customizable workflows are included. All workflows are supported by tutorials and Galaxy interactive tours to guide the users through the analyses step by step. ASaiM is implemented as Galaxy Docker flavour. It is scalable to many thousand datasets, but also can be used a normal PC. The associated source code is available under Apache 2 license at https://github.com/ASaiM/framework and documentation can be found online (http://asaim.readthedocs.io/)\n\nConclusionsBased on the Galaxy framework, ASaiM offers sophisticated analyses to scientists without command-line knowledge. ASaiM provides a powerful framework to easily and quickly explore microbiota data in a reproducible and transparent environment.

bioinformatics

Prediction of potential disease-associated microRNAs using structural perturbation method

MotivationThe identification of disease-related microRNAs(miRNAs) is an essential but challenging task in bioinformatics research. Similarity-based link prediction methods are often used to predict potential associations between miRNAs and diseases. In these methods, all unobserved associations are ranked by their similarity scores. Higher score indicates higher probability of existence. However, most previous studies mainly focus on designing advanced methods to improve the prediction accuracy while neglect to investigate the link predictability of the networks that present the miRNAs and diseases associations. In this work, we construct a bilayer network by integrating the miRNA-disease network, the miRNA similarity network and the disease similarity network. We use structural consistency as an indicator to estimate the link predictability of the related networks. On the basis of the indicator, a derivative algorithm, called structural perturbation method (SPM), is applied to predict potential associations between miRNAs and diseases.\n\nResultsThe link predictability of bilayer network is higher than that of miRNA-disease network, indicating that the prediction of potential miRNAs-diseases associations on bilayer network can achieve higher accuracy than based merely on the miRNA-disease network. A comparison between the SPM and other algorithms reveals the reliable performance of SPM which performed well in a 5-fold cross-validation. We test fifteen networks. The AUC values of SPM are higher than some well-known methods, indicating that SPM could serve as a useful computational method for improving the identification accuracy of miRNA-disease associations. Moreover, in a case study on breast neoplasm, 80% of the top-20 predicted miRNAs have been manually confirmed by previous experimental studies.\n\nAvailability and Implementationhttps://github.com/lecea/SPM-code.git\n\nContactlinyuan.lv@gmail.com, zouquan@nclab.net.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

GDCRNATools: an R/Bioconductor package for integrative analysis of lncRNA, miRNA, and mRNA data in GDC

The large-scale multidimensional omics data in the Genomic Data Commons (GDC) provides opportunities to investigate the crosstalk among different RNA species and their regulatory mechanisms in cancers. Easy-to-use bioinformatics pipelines are needed to facilitate such studies. We have developed a user-friendly R/Bioconductor package, named GDCRNATools, to facilitate downloading, organizing, and analyzing RNA data in GDC with an emphasis on deciphering the lncRNA-mRNA related competing endogenous RNAs (ceRNAs) regulatory network in cancers. Many widely used bioinformatics tools and databases are utilized in our package. Users can easily pack preferred downstream analysis pipelines or integrate their own pipelines into the workflow. Interactive shiny web apps built in GDCRNATools greatly improve visualization of results from the analysis.\n\nAvailabilityGDCRNATools is an R/Bioconductor package that is freely available at https://github.com/Jialab-UCR/GDCRNATools

bioinformatics

Prometheus: omics portals for interkingdom comparative genomic analyses

Functional analyses of genes are crucial for unveiling biological responses, for genetic engineering, and for developing new medicines. However, functional analyses have largely been restricted to model organisms, representing a major hurdle for functional studies and industrial applications. To resolve this, comparative genome analyses can be used to provide clues to gene functions as well as their evolutionary history. To this end, we present Prometheus (http://prometheus.kobic.re.kr),web-based omics portal that contains more than 17,215 sequences from prokaryotic and eukaryotic genomes. This portal supports interkingdom comparative analyses via a domain architecture-based gene identification system, Gene Search, and users can easily and rapidly identify single or entire gene sets in specific pathways. Bioinformatics tools for further analyses are provided in Prometheus or through BioExpress, a cloud-based bioinformatics analysis platform. Prometheus suggests a new paradigm for comparative analyses with large amounts of genomic information.

bioinformatics

Comparative Qualitative Phosphoproteomics Analysis Identifies Shared Phosphorylation Motifs and Associated Biological Processes in Flowering Plants

Phosphorylation is regarded as one of the most prevalent post-translational modifications and plays a key role in regulating cellular processes. In this work we carried out a comparative bioinformatics analysis of phosphoproteomics data, to profile two model species representing the largest subclasses in flowering plants the dicot Arabidopsis thaliana and the monocot Oryza sativa, to understand the extent to which phosphorylation signaling and function is conserved across evolutionary divergent plants. Using pre-existing mass spectrometry phosphoproteomics datasets and bioinformatic tools and resources, we identified 6,537 phosphopeptides from 3,189 phosphoproteins in Arabidopsis and 2,307 phosphopeptides from 1,613 phosphoproteins in rice. The relative abundance ratio of serine, threonine, and tyrosine phosphorylation sites in rice and Arabidopsis were highly similar: 88.3: 11.4: 0.4 and 86.7: 12.8: 0.5, respectively. Tyrosine phosphorylation shows features different from serine and threonine phosphorylation and was found to be more frequent in doubly-phosphorylated peptides in Arabidopsis. We identified phosphorylation sequence motifs in the two species to explore the similarities, finding nineteen pS motifs and two pT motifs that are shared in rice and Arabidopsis; among them are five novel motifs that have not previously been described in both species. The majority of shared motif-containing proteins were mapped to the same biological processes with similar patterns of fold enrichment, indicating high functional conservation. We also identified shared patterns of crosstalk between phosphoserines with motifs pSXpS, pSXXpS and pSXXXpS, where X is any amino acid, in both species indicating this is an evolutionary conserved signaling mechanism in flowering plants. However, our results are suggestive that there is greater co-occurrence of crosstalk between phosphorylation sites in Arabidopsis, and we were able to identify several pairs of motifs that are statistically significantly enriched to co-occur in Arabidopsis proteins, but not in rice.

bioinformatics

Asymptotically optimal minimizers schemes

MotivationThe minimizers technique is a method to sample k-mers that is used in many bioinformatics software to reduce computation, memory usage and run time. The number of applications using minimizers keeps on growing steadily. Despite its many uses, the theoretical understanding of minimizers is still very limited. In many applications, selecting as few k-mers as possible (i.e. having a low density) is beneficial. The density is highly dependent on the choice of the order on the k-mers. Different applications use different orders, but none of these orders are optimal. A better understanding of minimizers schemes, and the related local and forward schemes, will allow designing schemes with lower density, and thereby making existing and future bioinformatics tools even more efficient.\n\nResultsFrom the analysis of the asymptotic behavior of minimizers, forward and local schemes, we show that the previously believed lower bound on minimizers schemes does not hold, and that schemes with density lower than thought possible actually exist. The proof is constructive and leads to an efficient algorithm to compare k-mers. These orders are the first known orders that are asymptotically optimal. Additionally, we give improved bounds on the density achievable by the 3 type of schemes.\n\nContactgmarcais@cs.cmu.edu ckingsf@cs.cmu.edu

bioinformatics

Removing Contaminants from Metagenomic Databases

Metagenomic sequencing of patient samples is a very promising method for the diagnosis of human infections. Sequencing has the ability to capture all the DNA or RNA from pathogenic organisms in a human sample. However, complete and accurate characterization of the sequence, including identification of any pathogens, depends on the availability and quality of genomes for comparison. Thousands of genomes are now available, and as these numbers grow, the power of metagenomic sequencing for diagnosis should increase. However, recent studies have exposed the presence of contamination in published genomes, which when used for diagnosis increases the risk of falsely identifying the wrong pathogen.\n\nTo address this problem, we have developed a bioinformatics system for eliminating contamination as well as low-complexity genomic sequences in the draft genomes of eukaryotic pathogens. We applied this software to identify and remove human, bacterial, archaeal, and viral sequences present in a comprehensive database of all sequenced eukaryotic pathogen genomes. We also removed low-complexity genomic sequences, another source of false positives. Using this pipeline, we have produced a database of \"clean\" eukaryotic pathogen genomes for use with bioinformatics classification and analysis tools. We demonstrate that when attempting to find eukaryotic pathogens in metagenomic samples, the new database provides better sensitivity than one using the original genomes while offering a dramatic reduction in false positives.

bioinformatics

GTFtools: a Python package for analyzing various modes of gene models

SummaryGene-centric bioinformatics studies frequently involve calculation or extraction of various features of genes such as gene ID mapping, GC content calculation and different types of gene lengths, through manipulation of gene models that are often annotated in GTF format and available from ENSEMBL or GENCODE database. Such computation is essential for subsequent analysis such as intron retention detection where independent introns may need to be identified, converting RNA-seq read counts to FPKM where gene length is required, and obtaining flanking regions around transcription start sites. However, to our knowledge, a software package that is dedicated to analyzing various modes of gene models directly from GTF file is not publicly available. In this work, GTFtools (implemented in Python and not dependent on any non-python third-party software), a stand-alone command-line software that provides a set of functions to analyze various modes of gene models, is provided for facilitating routine bioinformatics studies where information about gene models needs to be calculated.\n\nAvailabilityGTFtools is freely available at www.genemine.org/gtftools.php\n\nContacthongdong@csu.edu.cn.

bioinformatics

MetaMap: An atlas of metatranscriptomic reads in human disease-related RNA-seq data

BackgroundWith the advent of the age of big data in bioinformatics, large volumes of data and high performance computing power enable researchers to perform re-analyses of publicly available datasets at an unprecedented scale. Ever more studies imply the microbiome in both normal human physiology and a wide range of diseases. RNA sequencing technology (RNA-seq) is commonly used to infer global eukaryotic gene expression patterns under defined conditions, including human disease-related contexts, but its generic nature also enables the detection of microbial and viral transcripts.\n\nFindingsWe developed a bioinformatic pipeline to screen existing human RNA-seq datasets for the presence of microbial and viral reads by re-inspecting the non-human-mapping read fraction. We validated this approach by recapitulating outcomes from 6 independent controlled infection experiments of cell line models and comparison with an alternative metatranscriptomic mapping strategy. We then applied the pipeline to close to 150 terabytes of publicly available raw RNA-seq data from >17,000 samples from >400 studies relevant to human disease using state-of-the-art high performance computing systems. The resulting data of this large-scale re-analysis are made available in the presented MetaMap resource.\n\nConclusionsOur results demonstrate that common human RNA-seq data, including those archived in public repositories, might contain valuable information to correlate microbial and viral detection patterns with diverse diseases. The presented MetaMap database thus provides a rich resource for hypothesis generation towards the role of the microbiome in human disease.

bioinformatics

VirTect: a computational method for detecting virus species from RNA-Seq and its application in head and neck squamous cell carcinoma

Next generation sequencing (NGS) provides an opportunity to detect viral species from RNA-seq data on human tissues, but existing computational approaches do not perform optimally on clinical samples. We developed a bioinformatics method called VirTect for detecting viruses in neoplastic human tissues using RNA-seq data. Here, we used VirTect to analyze RNA-seq data from 363 HNSCC (head and neck squamous cell carcinoma) patients and identified 22 HPV-induced HNSCCs. These predictions were validated by manual review of pathology reports on histopathologic specimens. Compared to two existing prediction methods, VirusFinder and VirusSeq, VirTect demonstrated superior performance with many fewer false positives and false negatives. The majority of HPV carcinogenesis studies thus far have been performed on cervical cancer and generalized to HNSCC. Our results suggest that HPV-induced HNSCC involves unique mechanisms of carcinogenesis, so understanding these molecular mechanisms will have a significant impact on therapeutic approaches and outcomes. In summary, VirTect can be an effective solution for the detection of viruses with NGS data, and can facilitate the clinicopathologic characterization of various types of cancers with broad applications for oncology.\n\nSignificance StatementWe developed a new bioinformatics tool, and reported the new inside of HPV carcinogenesis mechanism in HPV-induced head and neck squamous cell carcinoma (HNSCC). This novel bioin-formatics tool and the new knowledge of HPV-induced HNSCC will facilitate the development of target therapies for treating HNSCC.

bioinformatics

SoS Notebook: An Interactive Multi-Language Data Analysis Environment

MotivationComplex bioinformatic data analysis workflows involving multiple scripts in different languages can be difficult to consolidate, share, and reproduce. An environment that streamlines the entire processes of data collection, analysis, visualization and reporting of such multi-language analyses is currently lacking.\n\nResultsWe developed Script of Scripts (SoS) Notebook, a web-based notebook environment that allows the use of multiple scripting language in a single notebook, with data flowing freely within and across languages. SoS Notebook enables researchers to perform sophisticated bioinformatic analysis using the most suitable tools for different parts of the workflow, without the limitations of a particular language or complications of cross-language communications.\n\nAvailabilitySoS Notebook is hosted at http://vatlab.github.io/SoS/ and is distributed under a BSD license.\n\nContactbpeng@mdanderson.org

bioinformatics

An analysis of current state of the art software on nanopore metagenomic data

ContextOur insight into DNA is controlled through a process called sequencing. Until recently, it was only possible to sequence DNA into short strings called \"reads\". Nanopore is a new sequencing technology to produce significantly longer reads. Using nanopore sequencing, a single molecule of DNA can be sequenced without the need for time consuming PCR amplification (polymerase chain reaction is a technique used in molecular biology to amplify a single copy or a few copies of a segment of DNA across several orders of magnitude).\n\nAimsMetagenomics is the study of genetic material recovered from environmental samples. A research team from IBERS (Institute of Biological, Environmental & Rural Sciences) at Aberystwyth University have sampled metagenomes from a coal mine in South Wales using the Nanopore MinION and given initial taxonomic (classification of organisms) summaries of the contents of the microbial community.\n\nMethodsUsing various new software aimed for metagenomic data, we are interested to discover how well current bioinformatics software works with the data-set. We will conduct analysis and research into how well these new state of the art software works with this new long read data and try out some recent new developments for such analysis.\n\nResultsMost of the software we used worked very well: we gained understanding of the ACGT count and quality of the data. However some software for bioinformatics dont seem to work with nanopore data. Furthermore, we can conclude that low quality nanopore data may actually be quite average.

bioinformatics

DIMEdb: an integrated database and web service for metabolite identification in direct infusion mass spectrometery

MotivationMetabolomics involves the characterisation, identification, and quantification of small molecules (metabolites) that act as the reaction intermediates of biological processes. Over the past few years, we have seen wide scale improvements in data processing, database, and statistical analysis tools. Direct infusion mass spectrometery (DIMS) is a widely used platform that is able to produce a global fingerprint of the metabolome, without the requirement of a prior chromatographic step - making it ideal for wide scale high-throughput metabolomics analysis. In spite of these developments, metabolite identification still remains a key bottleneck in untargeted mass spectrometry-based metabolomics studies. The first step of the metabolite identification task is to query masses against a metaboite database to get putative metabolite annotations. Each existing metabolite database differs in a number of aspects including coverage, format, and accessibility - often limiting the user to a rudimentary web interface. Manually combining multiple search results for a single experiment where there may be potentially hundreds of masses to investigate becomes an incredibly arduous task.\n\nResultsTo facilitate unified access to metabolite information we have created the Direct Infusion MEtabolite database (DIMEdb), a comprehensive web-based metabolite database that contains over 80,000 metabolites sourced from a number of renowned metabolite databases of which can be utilised in the analysis and annotation of DIMS data. To demostrate the efficacy of DIMEdb, a simple use case for metabolic identification is presented. DIMEdb aims to provide a single point of access to metabolite information, and hopefully facilitate the development of much needed bioinformatic tools.\n\nAvailabilityDIMEdb is freely available at https://dimedb.ibers.aber.ac.uk.\n\nContactkeo7@aber.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Non-parametric and semi-parametric support estimation using SEquential RESampling random walks on biomolecular sequences

Non-parametric and semi-parametric resampling procedures are widely used to perform support estimation in computational biology and bioinformatics. Among the most widely used methods in this class is the standard bootstrap method, which consists of random sampling with replacement. While not requiring assumptions about any particular parametric model for resampling purposes, the bootstrap and related techniques assume that sites are independent and identically distributed (i.i.d.). The i.i.d. assumption can be an over-simplification for many problems in computational biology and bioinformatics. In particular, sequential dependence within biomolecular sequences is often an essential biological feature due to biochemical function, evolutionary processes such as recombination, and other factors.\n\nTo relax the simplifying i.i.d. assumption, we propose a new non-parametric/semi-parametric sequential resampling technique that generalizes \"Heads-or-Tails\" mirrored inputs, a simple but clever technique due to Landan and Graur. The generalized procedure takes the form of random walks along either aligned or unaligned biomolecular sequences. We refer to our new method as the SERES (or \"SEquential RESampling\") method.\n\nTo demonstrate the flexibility of the new technique, we apply SERES to two different applications - one involving aligned inputs and the other involving unaligned inputs. Using simulated and empirical data, we show that SERES-based support estimation yields comparable or typically better performance compared to state-of-the-art methods for both applications.

bioinformatics

Reproducible genomics analysis pipelines with GNU Guix

In bioinformatics, as well as other computationally-intensive research fields, there is a need for workflows that can reliably produce consistent output, independent of the software environment or configuration settings of the machine on which they are executed. Indeed, this is essential for controlled comparison between different observations or for the wider dissemination of workflows. Providing this type of reproducibility, however, is often complicated by the need to accommodate the myriad dependencies included in a larger body of software, each of which generally come in various versions. Moreover, in many fields (bioinformatics being a prime example), these versions are subject to continual change due to rapidly evolving technologies, further complicating problems related to reproducibility. Here, we propose a principled approach for building analysis pipelines and managing their dependencies. As a case study to demonstrate the utility of our approach, we present a set of highly reproducible pipelines for the analysis of RNA-seq, ChIP-seq, Bisulfite-seq, and single-cell RNA-seq. All pipelines process raw experimental data, and generate reports containing publication-ready plots and figures, with interactive report elements and standard observables. Users may install these highly reproducible packages and apply them to their own datasets without any special computational expertise beyond the use of the command line. We hope such a toolkit will provide immediate benefit to laboratory workers wishing to process their own data sets or bioinformaticians seeking to automate all, or parts of, their analyses. In the long term, we hope our approach to reproducibility will serve as a blueprint for reproducible workflows in other areas. Our pipelines, along with their corresponding documentation and sample reports, are available at http://bioinformatics.mdc-berlin.de/pigx

bioinformatics

HMP16SData: Efficient Access to the Human Microbiome Project through Bioconductor

Phase 1 of the NIH Human Microbiome Project (HMP) investigated 18 body subsites of 239 healthy American adults, to produce the first comprehensive reference for the composition and variation of the \"healthy\" human microbiome. Publicly-available data sets from amplicon sequencing of two 16S rRNA variable regions, with extensive controlled-access participant data, provide a reference for ongoing microbiome studies. However, utilization of these data sets can be hindered by the complex bioinformatic steps required to access, import, decrypt, and merge the various components in formats suitable for ecological and statistical analysis. The HMP16SData package provides count data for both 16S variable regions, integrated with phylogeny, taxonomy, public participant data, and controlled participant data for authorized researchers, using standard integrative Bioconductor data objects. By removing bioinformatic hurdles of data access and management, HMP16SData enables epidemiologists with only basic R skills to quickly analyze HMP data.

bioinformatics

PRESTO, a new tool for integrating large-scale -omics data and discovering disease-specific signatures

BackgroundCohesive visualization and interpretation of hyperdimensional, large-scale -omics data is an ongoing challenge, particularly for biologists and clinicians involved in current highly complex sequencing studies. Multivariate studies are often better suited towards non-linear network analysis than differential expression testing. Here, we present PRESTO, a PREdictive Stochastic neighbor embedding Tool for Omics, which allows unsupervised dimensionality reduction of multivariate data matrices with thousands of subjects or conditions. PRESTO is intuitively integrated into an interactive user interface that helps to visualize the multidimensional patterns in genome-wide transcriptomic data from basic science and clinical studies.\n\nResultsPRESTO was tested with multiple input omics platforms, including microarray and proteomics from both mouse and human clinical datasets. PRESTO can analyze up to tens of thousands of genes and shows no increase in processing time with a large number of samples or patients. In complex datasets, such as those with multiple time points, several patient groups, or diverse mouse strains, PRESTO outperformed conventional methods. Core co-expressed gene networks were intuitively grouped in clusters, or gates, after dimensionality reduction and remained consistent across users. Networks were identified and assigned to physiological and pathological functions that cannot be gleaned from conventional bioinformatics analyses. PRESTO detected gene networks from the natural variations among mouse macrophages and human blood leukocytes. We applied PRESTO to clinical transcriptomic and proteomic data from large patient cohorts and detected disease-defining signatures in antibody-mediated kidney transplant rejection, renal cell carcinoma, and relapsing acute myeloid leukemia (AML). In AML, PRESTO confirmed a previously described gene signature and found a new signature of 10 genes that is highly predictive of patient outcome.\n\nConclusionsPRESTO offers an important integration of powerful bioinformatics tools with an interactive user interface that increases data analysis accessibility beyond bioinformaticians and coders. Here, we show that PRESTO out performs conventional methods, such as DE analysis, in multi-dimensional datasets and can identify biologically relevant co-expression gene networks. In paired samples or time points, co-expression networks could be compared for insight into longitudinal regulatory mechanisms. Additionally, PRESTO identified disease-specific signatures in clinical datasets with highly significant diagnostic and prognostic potential.

bioinformatics

CRISPRCloud2: A cloud-based platform for deconvoluting CRISPR screen data

The simplicity and cost-effectiveness of CRISPR technology have made high-throughput pooled screening approaches available to many. However, the large amount of sequencing data derived from these studies yields often unwieldy datasets requiring considerable bioinformatic resources to deconvolute data; a feature which is simply not accessible to many wet labs. To address these needs, we have developed a cloud-based webtool CRISPRCloud2 that provides a state-of-the-art accuracy in mapping short reads to CRISPR library, a powerful statistical test that aggregates information across multiple sgRNAs targeting the same gene, a user-friendly data visualization and query interface, as well as easy linking to other CRISPR tools and bioinformatics resources for target prioritization. CRISPRCloud2 is a one-stop shop for labs analyzing CRISPR screen data.

bioinformatics