Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

dBBQs : dataBase of Bacterial Quality scores

BackgroundIt is well-known that genome sequencing technologies are becoming significantly cheaper and faster. As a result of this, the exponential growth in sequencing data in public databases allows us to explore ever growing large collections of genome sequences. However, it is less known that the majority of available sequenced genome sequences in public databases are not complete, drafts of varying qualities. We have calculated quality scores for around 100,000 bacterial genomes from all major genome repositories and put them in a fast and easy-to-use database.\n\nResultsProkaryotic genomic data from all sources were collected and combined to make a non-redundant set of bacterial genomes. The genome quality score for each was calculated by four different measurements: assembly quality, number of rRNA and tRNA genes, and the occurrence of conserved functional domains. The dataBase of Bacterial Quality scores (dBBQs) was designed to store and retrieve quality scores. It offers fast searching and download features which the result can be used for further analysis. In addition, the search results are shown in interactive JavaScript chart framework using DC.js. The analysis of quality scores across major public genome databases find that around 68% of the genomes are of acceptable quality for many uses.\n\nConclusionsdBBQs (available at http://arc-gem.uams.edu/dbbqs) provides genome quality scores for all available prokaryotic genome sequences with a user-friendly Web-interface. These scores can be used as cut-offs to get a high-quality set of genomes for testing bioinformatics tools or improving the analysis. Moreover, all data of the four measurements that were combined to make the quality score for each genome, which can potentially be used for further analysis. dBBQs will be updated regularly and is freely use for non-commercial purpose.

bioinformatics

SnapperDB: A database solution for routine sequencing analysis of bacterial isolates

Real-time surveillance of infectious disease using whole genome sequencing data poses challenges in both result generation and communication. SnapperDB represents a set of tools to store bacterial variant data and facilitate reproducible and scalable analysis of bacterial populations. We also introduce the SNP address nomenclature to describe the relationship between isolates in a population to the single nucleotide resolution.\n\nSummaryWe announce the release of SnapperDB v1.0 a program for scalable routine SNP analysis and storage of microbial populations.\n\nAvailabilitySnapperDB is implemented as a python application under the open source BSD license. All code and user guides are available at https://github.com/phe-bioinformatics/snapperdb.\n\nContacttim.dallman@phe.gov.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Efficient management and analysis of large-scale genome-wide data with two R packages: bigstatsr and bigsnpr

MotivationGenome-wide datasets produced for association studies have dramatically increased in size over the past few years, with modern datasets commonly including millions of variants measured in dozens of thousands of individuals. This increase in data size is a major challenge severely slowing down genomic analyses. Specialized software for every part of the analysis pipeline have been developed to handle large genomic data. However, combining all these software into a single data analysis pipeline might be technically difficult.\n\nResultsHere we present two R packages, bigstatsr and bigsnpr, allowing for management and analysis of large scale genomic data to be performed within a single comprehensive framework. To address large data size, the packages use memory-mapping for accessing data matrices stored on disk instead of in RAM. To perform data pre-processing and data analysis, the packages integrate most of the tools that are commonly used, either through transparent system calls to existing software, or through updated or improved implementation of existing methods. In particular, the packages implement a fast derivation of Principal Component Analysis, functions to remove SNPs in Linkage Disequilibrium, and algorithms to learn Polygenic Risk Scores on millions of SNPs. We illustrate applications of the two R packages by analysing a case-control genomic dataset for the celiac disease, performing an association study and computing Polygenic Risk Scores. Finally, we demonstrate the scalability of the R packages by analyzing a simulated genome-wide dataset including 500,000 individuals and 1 million markers on a single desktop computer.\n\nAvailabilityhttps://privefl.github.io/bigstatsr/ & https://privefl.github.io/bigsnpr/\n\nContactflorian.prive@univ-grenoble-alpes.fr & michael.blum@univ-grenoble-alpes.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Extracting Evidence Fragments for Distant Supervision of Molecular Interactions

Abstract.We describe a methodology for automatically extracting evidence fragments from a set of biomedical experimental research articles. These fragments provide the primary description of evidence that is presented in the papers figures. They elucidate the goals, methods, results and interpretations of experiments that support the original scientific contributions the study being reported. Within this paper, we describe our methodology and showcase an example data set based on the European Bioinformatics Institutes INTACT database (http://www.ebi.ac.uk/intact/). Using figure codes as anchors, we linked evidence fragments to INTACT data records as an example of distant supervision so that we could use INTACTs preexisting, manually-curated structured interaction data to act as a gold standard for machine reading experiments. We report preliminary baseline event extraction measures from this collection based on a publicly available, machine reading system (REACH). We use semantic web standards for our data and provide open access to all source code.

bioinformatics

DAFi: A Directed Recursive Filtering and Clustering Approach to Data-Driven Identification of Cell Populations from Polychromatic Flow Cytometry Data

Computational methods for identification of cell populations from high-dimensional flow cytometry data are changing the paradigm of cytometry bioinformatics. Data clustering is the most common computational approach to unsupervised identification of cell populations from multidimensional cytometry data. We found that combining recursive filtering and clustering with constraints converted from the user manual gating strategy can effectively identify overlapping and rare cell populations from smeared data that would have been difficult to resolve by either a single run of data clustering or manual segregation. We named this new method DAFi: Directed Automated Filtering and Identification of cell populations. Design of DAFi preserves the data-driven characteristics of unsupervised clustering for identifying novel cell-based biomarkers, but also makes the results interpretable to experimental scientists as in supervised classification through mapping and merging the high-dimensional data clusters into the user-defined 2D gating hierarchy. By recursive data filtering before clustering, DAFi can uncover small local clusters which are otherwise difficult to identify due to the statistical interference of the irrelevant major clusters. Quantitative assessment of cell type specific characteristics demonstrates that the population proportions calculated by DAFi, while being highly consistent with those by expert centralized manual gating, have smaller technical variance than those from individual manual gating analysis. Visual examination of the dot plots showed that the boundaries of the DAFi-identified cell populations followed the natural shapes of the data distributions. To further exemplify the utility of DAFi, we show that DAFi can incorporate the FLOCK clustering method to identify novel cell-based biomarkers. Implementation of DAFi supports options including clustering, bisecting, slope-based gating, and reversed filtering to meet various auto-gating needs from different scientific use cases.

bioinformatics

Gauss-power mixing distributions comprehensively describe stochastic variations in RNA-seq data

MotivationGene expression levels exhibit stochastic variations among genetically identical organisms under the same environmental conditions. In many recent transcriptome analyses based on RNA sequencing (RNA-seq), variations in gene expression levels among replicates were assumed to follow a negative binomial distribution although the physiological basis of this assumption remain unclear.\n\nResultsIn this study, RNA-seq data were obtained from Arabidopsis thaliana under eight conditions (21-27 replicates), and the characteristics of gene-dependent distribution profiles of gene expression levels were analyzed. For A. thaliana and Saccharomyces cerevisiae, the distribution profiles could be described by a Gauss-power mixing distribution derived from a simple model of a stochastic transcriptional network containing a feedback loop. The distribution profiles of gene expression levels were roughly classified as Gaussian, power law-like containing a long tail, and mixed. The fitting function predicted that gene expression levels with long-tailed distributions would be strongly influenced by feedback regulation. Thus, the features of gene expression levels are correlated with their functions, with the levels of essential genes tending to follow a Gaussian distribution and those of genes encoding nucleic acid-binding proteins and transcription factors exhibiting long-tailed distributions.\n\nAvailabilityFastq files of RNA-seq experiments were deposited into the DNA Data Bank of Japan Sequence Read Archive as accession no. DRA005887. Quantified expression data are available in supplementary information.\n\nContactawa@hiroshima-u.ac.jp\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Integrative pipeline for profiling DNA copy number and inferring tumor phylogeny

SummaryCopy number variation is an important and abundant source of variation in the human genome, which has been associated with a number of diseases, especially cancer. Massively parallel next-generation sequencing allows copy number profiling with fine resolution. Such efforts, however, have met with mixed successes, with setbacks arising partly from the lack of reliable analytical methods to meet the diverse and unique challenges arising from the myriad experimental designs and study goals in genetic studies. In cancer genomics, detection of somatic copy number changes and profiling of allele-specific copy number (ASCN) are complicated by experimental biases and artifacts as well as normal cell contamination and cancer subclone admixture. Furthermore, careful statistical modeling is warranted to reconstruct tumor phylogeny by both somatic ASCN changes and single nucleotide variants. Here we describe a flexible computational pipeline, MARATHON, which integrates multiple related statistical software for copy number profiling and downstream analyses in disease genetic studies.\n\nAvailability and implementationMARATHON is publicly available at https://github.com/yuchaojiang/MARATHON.\n\nContactyuchaoj@email.unc.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Haystack: systematic analysis of the variation of epigenetic states and cell-type specific regulatory elements

MotivationWith the increasing amount of genomic and epigenomic data in the public domain, a pressing challenge is how to integrate these data to investigate the role of epigenetic mechanisms in regulating gene expression and maintenance of cell-identity. To this end, we have implemented a computational pipeline to systematically study epigenetic variability and uncover regulatory DNA sequences that play a role in gene regulation.\n\nResultsHaystack is a bioinformatics pipeline to characterize hotspots of epigenetic variability across different cell-types as well as cell-type specific cis-regulatory elements along with their corresponding transcription factors. Our approach is generally applicable to any epigenetic mark and provides an important tool to investigate cell-type identity and the mechanisms underlying epigenetic switches during development. Additionally, we make available a set of precomputed tracks for a number of epigenetic marks across several cell types. These precomputed results may be used as an independent resource for functional annotation of the human genome.\n\nAvailabilityThe Haystack pipeline is implemented as an open-source, multiplatform, Python package called haystack_bio available at https://github.com/pinellolab/haystack_bio.\n\nContactlpinello@mgh.harvard.edu, gcyuan@jimmy.harvard.edu

bioinformatics

Delta integrates 3D physical structure with topology and genomic data of chromosomes

MotivationThe regulation of gene transcription and DNA replication are tightly associated with the 3D chromosomal structures and genomic features, e.g. epigenetic marks, transcription factor bindings and non-coding RNAs. The interaction between the features and the chromosomal structures forming a multilayer 3D regulatory network. Therefore, it is necessary to integrate the physical 3D architecture of genome and features to comprehensive depict their connection to gene regulation.\n\nResultsHere, we present an integrative visualization and analysis platform, Delta, to facilitate visually annotating and exploring the 3D physical architecture of genomes. Delta takes Hi-C or ChIA-PET contact matrix as input and predicts the topology associated domains and chromatin loops in the genome, and generates a physical 3D model which represents the plausible consensus 3D structure of the genome. Delta features a highly interactive visualization tool, which enhanced the integration of genome topology/physical structure and extensive genome annotation, by juxtaposition of the 3D model with diverse genomic assay outputs. Finally, we showcased that Delta could be helpful to reveal potentially interesting findings by a case study on the {beta}-globin gene region.\n\nAvailability and implementationhttp://delta.big.ac.cn/.\n\nContacttangbx@big.ac.cn.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Practical computational reproducibility in the life sciences

Many areas of research suffer from poor reproducibility. This problem is particularly acute in computationally intensive domains where results rely on a series of complex methodological decisions that are not well captured by traditional publication approaches. Various guidelines have emerged for achieving reproducibility, but practical implementation of these practices remains difficult. This is because reproducing published computational analyses requires installing many software tools plus associated libraries, connecting tools together into the complete pipeline, and specifying parameters. Here we present a suite of recently emerged technologies which make computational reproducibility not just possible, but, finally, practical in both time and effort. By combining a system for building highly portable packages of bioinformatics software, containerization and virtualization technologies for isolating reusable execution environments for these packages, and an integrated workflow system that automatically orchestrates the composition of these packages for entire pipelines, an unprecedented level of computational reproducibility can be achieved.

bioinformatics

Identifying accurate metagenome and amplicon software via a meta-analysis of benchmarking studies

Environmental DNA sequencing has rapidly become a widely-used technique for investigating a range of questions, particularly related to health and environmental monitoring. There has also been a proliferation of bioinformatic tools for analysing metagenomic and amplicon datasets, which makes selecting adequate tools a significant challenge. A number of benchmark studies have been undertaken; however, these can present conflicting results. We have applied a robust Z-score ranking procedure and a network meta-analysis method to identify software tools that are generally accurate for mapping DNA sequences to taxonomic hierarchies. Based upon these results we have identified some tools and computational strategies that produce robust predictions.

bioinformatics

Bioconda: A sustainable and comprehensive software distribution for the life sciences

We present Bioconda (https://bioconda.github.io), a distribution of bioinformatics software for the lightweight, multiplatform and language-agnostic package manager Conda. Currently, Bioconda offers a collection of over 3000 software packages, which is continuously maintained, updated, and extended by a growing global community of more than 200 contributors. Bioconda improves analysis reproducibility by allowing users to define isolated environments with defined software versions, all of which are easily installed and managed without administrative privileges.

bioinformatics

So you think you can PLS-DA?

BackgroundPartial Least-Squares Discriminant Analysis (PLS-DA) is a popular machine learning tool that is gaining increasing attention as a useful feature selector and classifier. In an effort to understand its strengths and weaknesses, we performed a series of experiments with synthetic data and compared its performance to its close relative from which it was initially invented, namely Principal Component Analysis (PCA). ResultsWe demonstrate that even though PCA ignores the information regarding the class labels of the samples, this unsupervised tool can be remarkably effective as a feature selector. In some cases, it outperforms PLS-DA, which is made aware of the class labels in its input. Our experiments range from looking at the signal-to-noise ratio in the feature selection task, to considering many practical distributions and models encountered when analyzing bioinformatics and clinical data. Other methods were also evaluated. Finally, we analyzed an interesting data set from 396 vaginal microbiome samples where the ground truth for the feature selection was available. All the 3D figures shown in this paper as well as the supplementary ones can be viewed interactively at http://biorg.cs.fiu.edu/plsda ConclusionsOur results highlighted the strengths and weaknesses of PLS-DA in comparison with PCA for different underlying data models.

bioinformatics

MeShClust: an intelligent tool for clustering DNA sequences

Sequence clustering is a fundamental step in analyzing DNA sequences. Widely-used software tools for sequence clustering utilize greedy approaches that are not guaranteed to produce the best results. These tools are sensitive to one parameter that determines the similarity among sequences in a cluster. Often times, a biologist may not know the exact sequence similarity. Therefore, clusters produced by these tools do not likely match the real clusters comprising the data if the provided parameter is inaccurate. To overcome this limitation, we adapted the mean shift algorithm, an unsupervised machine-learning algorithm, which has been used successfully thousands of times in fields such as image processing and computer vision. The theory behind the mean shift algorithm, unlike the greedy approaches, guarantees convergence to the modes, e.g. cluster centers. Here we describe the first application of the mean shift algorithm to clustering DNA sequences. MeShClust is one of few applications of the mean shift algorithm in bioinformatics. Further, we applied supervised machine learning to predict the identity score produced by global alignment using alignment-free methods. We demonstrate MeShClusts ability to cluster DNA sequences with high accuracy even when the sequence similarity parameter provided by the user is not very accurate.

bioinformatics

Identification and prioritisation of causal variants in human genetic disorders from exome or whole genome sequencing data

With genome sequencing entering the clinics as diagnostic tool to study genetic disorders, there is an increasing need for bioinformatics solutions that enable precise causal variant identification in a timely manner.\n\nBackgroundWorkflows for the identification of candidate disease-causing variants perform usually the following tasks: i) identification of variants; ii) filtering of variants to remove polymorphisms and technical artifacts; and iii) prioritization of the remaining variants to provide a small set of candidates for further analysis.\n\nMethodsHere, we present a pipeline designed to identify variants and prioritize the variants and genes from trio sequencing or pedigree-based sequencing data into different tiers.\n\nResultsWe show how this pipeline was applied in a study of patients with neurodevelopmental disorders of unknown cause, where it helped to identify the causal variants in more than 35% of the cases.\n\nConclusionsClassification and prioritization of variants into different tiers helps to select a small set of variants for downstream analysis.

bioinformatics

Non-biological synthetic spike-in controls and the AMPtk software pipeline improve fungal high throughput amplicon sequencing data

High throughput amplicon sequencing (HTAS) of conserved DNA regions is a powerful technique to characterize microbial communities. Recently, spike-in mock communities have been used to measure accuracy of sequencing platforms and data analysis pipelines. To assess the ability of sequencing platforms and data processing pipelines using fungal ITS amplicons, we created two ITS spike-in control mock communities composed of cloned DNA in plasmids: a biological mock community (BioMock), consisting of ITS sequences from fungal taxa, and a synthetic mock community (SynMock), consisting of non-biological ITS-like sequences. Using these spike-in controls we show that: 1) a non-biological synthetic control (e.g., SynMock) is the best solution for parameterizing bioinformatics pipelines, 2) pre-clustering steps for variable length amplicons are critically important, 3) a major source of bias is attributed to initial PCR reactions and thus HTAS read abundances are typically not representative of starting values. We developed AMPtk, a versatile software solution equipped to deal with variable length amplicons and quality filter HTAS data based on spike-in controls. While we describe herein a non-biological synthetic mock community for ITS sequences, the concept and AMPtk software can be widely applied to any HTAS dataset to improve data quality.\n\nAvailability and Implementation - AMPtk is publically available at https://github.com/nextgenusfs/amptk. All primary data and data analysis done in this manuscript are available via the Open Science Framework (https://osf.io/4xd9r/). The SynMock sequences and the script to produce them are available in the OSF repository ((https://osf.io/4xd9r/) as well as packaged into AMPtk distributions.

bioinformatics

SAFE-clustering: Single-cell Aggregated (From Ensemble) Clustering for Single-cell RNA-seq Data

MotivationAccurately clustering cell types from a mass of heterogeneous cells is a crucial first step for the analysis of single-cell RNA-seq (scRNA-Seq) data. Although several methods have been recently developed, they utilize different characteristics of data and yield varying results in terms of both the number of clusters and actual cluster assignments.\n\nResultsHere, we present SAFE-clustering, Single-cell Aggregated (From Ensemble) clustering, a flexible, accurate and robust method for clustering scRNA-Seq data. SAFE-clustering takes as input, results from multiple clustering methods, to build one consensus solution. SAFE-clustering currently embeds four state-of-the-art methods, SC3, CIDR, Seurat and t-SNE + k-means; and ensembles solutions from these four methods using three hypergraph-based partitioning algorithms. Extensive assessment across 12 datasets with the number of clusters ranging from 3 to 14, and the number of single cells ranging from 49 to 32,695 showcases the advantages of SAFE-clustering in terms of both cluster number (18.9 - 50.0% reduction in absolute deviation to the truth) and cluster assignment (on average 28.9% improvement, and up to 34.5% over the best of the four methods, measured by adjusted rand index). Moreover, SAFE-clustering is computationally efficient to accommodate large datasets, taking <10 minutes to process 28,733 cells.\n\nAvailability and implementationSAFE-clustering, including source codes and tutorial, is free available on the web at http://yunliweb.its.unc.edu/safe/.\n\nContactyunli@med.unc.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Biomarker identification for statin sensitivity of cancer cell lines

Statins are potent cholesterol reducing drugs that have been shown to reduce tumor cell proliferation in vitro and tumor growth in animal models. Moreover, retrospective human cohort studies demon-strated decreased cancer-specific mortality in patients taking statins. We previously implicated membrane E-cadherin expression as both a marker and mechanism for resistance to atorvastatin-mediated growth suppression of cancer cells; however, a transcriptome-profile-based biomarker signature for statin sensitivity has not yet been reported. Here, we utilized transcriptome data from fourteen NCI-60 cancer cell lines and their statin dose-response data to produce gene expression signatures that identify statin sensitive and resistant cell lines. We experimentally confirmed the validity of the identified biomarker signature in an independent set of cell lines and extended this signature to generate a proposed statin-sensitive subset of tumors listed in the TCGA database. Finally, we predicted drugs that would synergize with statins and found several predicted combination therapies to be experimentally confirmed. The combined bioinformatics-experimental approach described here can be used to generate an initial biomarker sensitivity for statin therapy.

bioinformatics