Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

LoRTE: Detecting transposon-induced genomic variants using low coverage PacBio long read sequences

MotivationPopulation genomic analysis of transposable elements has greatly benefited from recent advances of sequencing technologies. However, the propensity of transposable elements to nest in highly repeated regions of genomes limits the efficiency of bioinformatic tools when short read sequences technology is used.\n\nResultsLoRTE is the first tool able to use PacBio long read sequences to identify transposon deletions and insertions between a reference genome and genomes of different strains or populations. Tested against Drosophila melanogaster PacBio datasets, LoRTE appears to be a reliable and broadly applicable tools to study the dynamic and evolutionary impact of transposable elements using low coverage, long read sequences.\n\nAvailability and ImplementationLoRTE is available at http://www.egce.cnrs-gif.fr/?p=6422. It is written in Python 2.7 and only requires the NCBI BLAST + package. LoRTE can be used on standard computer with limited RAM resources and reasonable running time even with large datasets.\n\nContactjonathan.filee@ecge.cnrs-gif.fr

Bioinformatics

SVScore: An Impact Prediction Tool For Structural Variation

MotivationStructural variation (SV) is an important and diverse source of human genome variation. Over the past several years, much progress has been made in the area of SV detection, but predicting the functional impact of SVs discovered in whole genome sequencing (WGS) studies remains extremely challenging. Accurate SV impact prediction is especially important for WGS-based rare variant association studies and studies of rare disease.\n\nResultsHere we present SVScore, a computational tool for in silico SV impact prediction. SVScore aggregates existing per-base single nucleotide polymorphism pathogenicity scores across relevant genomic intervals for each SV in a manner that considers variant type, gene features, and uncertainty in breakpoint location. We show that in a Finnish cohort, the allele frequency spectrum of SVs with high impact scores is strongly skewed toward lower frequencies, suggesting that these variants are under purifying selection. We further show that SVScore identifies deleterious variants more effectively than naive alternative methods. Finally, our results indicate that high-scoring tandem duplications may be under surprisingly strong selection relative to high-scoring deletions, suggesting that duplications may be more deleterious than previously thought. In conclusion, SVScore provides pathogenicity prediction for SVs that is both informative and meaningful for understanding their functional role in disease.\n\nAvailabilitySVScore is implemented in Perl and available freely at {{http://www.github.com/lganel/SVScore}} for use under the MIT license.\n\nContactihall@wustl.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

MTQuant: "Seeing" Beyond the Diffraction Limit in Fluorescence Images to Quantify Neuronal Microtubule Organization

MotivationMicrotubules (MTs) are polarized polymers that are critical for cell structure and axonal transport. They form a bundle in neurons, but beyond that, their organization is relatively unstudied.\n\nResultsWe present MTQuant, a method for quantifying MT organization using light microscopy, which distills three parameters from MT images: the spacing of MT minus-ends, their average length, and the average number of MTs in a cross-section of the bundle. This method allows for robust and rapid in vivo analysis of MTs, rendering it more practical and more widely applicable than commonly-used electron microscopy reconstructions. MTQuant was successfully validated with three ground truth data sets and applied to over 3000 images of MTs in a C. elegans motor neuron.\n\nAvailabilityMATLAB code is available at http://roscoope.github.io/MTQuant\n\nContacthorowitz@stanford.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Annotation and differential analysis of alternative splicing using de novo assembly of RNAseq data

Genome-wide analyses reveal that more than 90% of multi exonic human genes produce at least two transcripts through alternative splicing (AS). Various bioinformatics methods are available to analyze AS from RNAseq data. Most methods start by mapping the reads to an annotated reference genome, but some start by a de novo assembly of the reads. In this paper, we present a systematic comparison of a mapping-first approach (FO_SCPLOWAC_SCPLOWRLO_SCPLOWINEC_SCPLOW) and an assembly-first approach (KO_SCPLOWISC_SCPLOWSO_SCPLOWPLICEC_SCPLOW). These two approaches are event-based, as they focus on the regions of the transcripts that vary in their exon content. We applied these methods to an RNAseq dataset from a neuroblastoma SK-N-SH cell line (ENCODE) differentiated or not using retinoic acid. We found that the predictions of the two pipelines overlapped (70% of exon skipping events were common), but with noticeable differences. The assembly-first approach allowed to find more novel variants, including novel unannotated exons and splice sites. It also predicted AS in families of paralog genes. The mapping-first approach allowed to find more lowly expressed splicing variants, and was better in predicting exons overlapping repeated elements. This work demonstrates that annotating AS with a single approach leads to missing a large number of candidates. We further show that these candidates cannot be neglected, since many of them are differentially regulated across conditions, and can be validated experimentally. We therefore advocate for the combine use of both mapping-first and assembly-first approaches for the annotation and differential analysis of AS from RNAseq data.

Bioinformatics

Phylo-Node: a molecular phylogenetic toolkit using Node.js

Background: Node.js is an open-source and cross-platform environment that provides a JavaScript codebase for back-end server-side applications. JavaScript has been used to develop very fast, and user-friendly front-end tools for bioinformatic and phylogenetic analyses. However, no such toolkits are available using Node.js to conduct comprehensive molecular phylogenetic analysis.\n\nResults: To address this problem, I have developed, Phylo-Node, which was developed using Node.js and provides a stable and scalable toolkit that allows the user to go from sequence retrieval to phylogeny reconstruction. Phylo-Node can execute the analysis and process the resulting outputs from a suite of software options that provides tools for sequence retrieval, alignment, primer design, evolutionary modeling, and phylogeny reconstruction. Furthermore, Phylo-Node enables the user to deploy server dependent applications, and also provides simple integration and interoperation with other Node modules and languages using Node inheritance patterns, and a customized piping module to support the production of diverse pipelines.\n\nConclusions: Phylo-Node is open-source and freely available to all users without sign-up or login requirements. All source code and user guidelines are openly available at the GitHub repository: https://github.com/dohalloran/Phylo-Node

Bioinformatics

Identification of outcome-related driver mutations in cancer using conditional co-occurrence distributions

The methods proposed for the detection of cancer driver mutations are based on the estimation of background mutation rate, impact on protein function, or network influence. Instead, we focus on those influencing patient survival. For this, an approximation of the log-rank test has been systematically applied even though it assumes a large and similar number of patients in both risk groups, which is violated in cancer genomics. Here, we propose VALORATE, a novel algorithm for the estimation of the null distribution for the log-rank test independently of the number of mutations. VALORATE is based on conditional distributions of the co-occurrences between events and mutations. The results using simulations, comparisons with other methods, TCGA and ICGC cancer datasets, and validations, suggests that VALORATE is accurate, fast, and can identify known and novel gene mutations. Our proposal and results may have important implications in cancer biology, in bioinformatics analyses, and ultimately in precision medicine.

Bioinformatics

GenomeScope: Fast reference-free genome profiling from short reads

SummaryGenomeScope is an open-source web tool to rapidly estimate the overall characteristics of a genome, including genome size, heterozygosity rate, and repeat content from unprocessed short reads. These features are essential for studying genome evolution, and help to choose parameters for downstream analysis. We demonstrate its accuracy on 324 simulated and 16 real datasets with a wide range in genome sizes, heterozygosity levels, and error rates.\n\nAvailability and Implementationhttp://genomescope.org, https://github.com/schatzlab/genomescope.git\n\nContactmschatz@jhu.edu.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Accurate contact predictions for thousands of protein families using PconsC3

Protein structure prediction was for decades one of the grand unsolved challenges in bioinformatics. A few years ago it was shown that by using a maximum entropy approach to describe couplings between columns in a multiple sequence alignment it was possible to significantly increase the accuracy of residue contact predictions. For very large protein families with more than 1000 effective sequences the accuracy is sufficient to produce accurate models of proteins as well as complexes. Today, for about half of all Pfam domain families no structure is known, but unfortunately most of these families have at most a few hundred members, i.e. are too small for existing contact prediction methods. To extend accurate contact predictions to the thousands of smaller protein families we present PconsC3, an improved method for protein contact predictions that can be used for families with as little as 100 effective sequence members. We estimate that PconsC3 provides accurate contact predictions for up to 4646 Pfam domain families. In addition, PconsC3 outperforms previous methods significantly independent on family size, secondary structure content, contact range, or the number of selected contacts. This improvement translates into improved de-novo prediction of three-dimensional structures. PconsC3 is available as a web server and downloadable version at http://c3.pcons.net. The downloadable version is free for all to use and licensed under the GNU General Public License, version 2.

Bioinformatics

Using predictive specificity to determine when gene set analysis is biologically meaningful

Gene set analysis, which translates gene lists into enriched functions, is among the most common bioinformatic methods. Yet few would advocate taking the results at face value. Not only is there no agreement on the algorithms themselves, there is no agreement on how to benchmark them. In this paper, we evaluate the robustness and uniqueness of enrichment results as a means of assessing methods even where correctness is unknown. We show that heavily annotated (\"multifunctional\") genes are likely to appear in genomics study results and drive the generation of biologically non-specific enrichment results as well as highly fragile significances. By providing a means of determining where enrichment analyses report non-specific and non-robust findings, we are able to assess where we can be confident in their use. We find significant progress in recent bias correction methods for enrichment and provide our own software implementation. Our approach can be readily adapted to any pre-existing package.

Bioinformatics

deSPI: efficient classification of metagenomic reads with lightweight de Bruijn graph-based reference indexing

SummaryIn metagenomic studies, fast and effective tools are on wide demand to implement taxonomy classification for upto billions of reads. Herein, we propose deSPI, a novel read classification method that classifies reads by recognizing and analyzing the matches between reads and reference with de Bruijn graph-based lightweight reference indexing. deSPI has faster speed with relatively small memory footprint, meanwhile, it can also achieve higher or similar sensitivity and accuracy.\n\nAvailabilitythe C++ source code of deSPI is available at https://github.com/hitbc/deSPI\n\nContactydwang@hit.edu.cn\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

Ontology-based workflow extraction from texts using word sense disambiguation

This paper introduces a method for automatic workflow extraction from texts using Process-Oriented Case-Based Reasoning (POCBR). While the current workflow management systems implement mostly different complicated graphical tasks based on advanced distributed solutions (e.g. cloud computing and grid computation), workflow knowledge acquisition from texts using case-based reasoning represents more expressive and semantic cases representations. We propose in this context, an ontology-based workflow extraction framework to acquire processual knowledge from texts. Our methodology extends classic NLP techniques to extract and disambiguate tasks in texts. Using a graph-based representation of workflows and a domain ontology, our extraction process uses a context-based approach to recognize workflow components : data and control flows. We applied our framework in a technical domain in bioinformatics : i.e. phylogenetic analyses. An evaluation based on workflow semantic similarities on a gold standard proves that our approach provides promising results in the process extraction domain. Both data and implementation of our framework are available in : http://labo.bioinfo.uqam.ca/tgrowler.

bioinformatics

Pavian: Interactive analysis of metagenomics data for microbiomics and pathogen identification

SummaryPavian is a web application for exploring metagenomics classification results, with a special focus on infectious disease diagnosis. Pinpointing pathogens in metagenomics classification results is often complicated by host and laboratory contaminants as well as many non-pathogenic microbiota. With Pavian, researchers can analyze, display and transform results from the Kraken and Centrifuge classifiers using interactive tables, heatmaps and flow diagrams. Pavian also provides an alignment viewer for validation of matches to a particular genome.\n\nAvailability and implementationPavian is implemented in the R language and based on the Shiny framework. It can be hosted on Windows, Mac OS X and Linux systems, and used with any contemporary web browser. It is freely available under a GPL-3 license from http://github.com/fbreitwieser/pavian. Furthermore a Docker image is provided at https://hub.docker.com/r/florianbw/pavian.\n\nContactfbreitw1@jhu.edu\n\nSupplementary informationSupplementary data is available at Bioinformatics online.

bioinformatics

Assessment of Antibody Library Diversity through Next Generation Sequencing and Technical Error Compensation.

Antibody libraries are important resources to derive antibodies to be used for a wide range of applications, from structural and functional studies to intracellular protein interference studies to developing new diagnostics and therapeutics. Whatever the goal, the key parameter for an antibody library is its diversity, i.e. the number of distinct elements in the collection, which directly reflects the probability of finding in the library an antibody against a given antigen, of sufficiently high affinity. Quantitative evaluation of antibody library diversity and quality has been for a long time inadequately addressed, due to the high similarity and length of the sequences of the library. Diversity was usually inferred by the transformation efficiency and tested either by fingerprinting and/or sequencing of a few hundred random library elements. Inferring diversity from such a small sample is, however, very rudimental and gives limited information about the real complexity, because complexity does not scale linearly with sample size. Next-generation sequencing (NGS) has opened new ways to tackle the antibody library diversity quality assessment. However, much remains to be done to fully exploit the potential of NGS for the quantitative analysis of antibody repertoires and to overcome current limitations. To obtain a more reliable antibody library complexity estimate here we show a new, PCR-free, NGS approach to sequence antibody libraries on Illumina platform, coupled to a new bioinformatic analysis and software (Diversity Estimator of Antibody Library, DEAL) that allows to reliably estimate the diversity, taking in consideration the sequencing error.

bioinformatics

GLASS: assisted and standardized assessment of gene variations from Sanger sequence trace data

MotivationSanger sequencing remains the reference method for sequence variant detection, especially in a clinical setting. However, chromatogram interpretation often requires manual inspection and in some cases considerable expertise. Additionally, variant reporting and nomenclature is typically left to the user, which can lead to inconsistencies.\n\nResultsWe introduce GLASS, a tool built to assist with the assessment of gene variations in Sanger sequencing data. Critically, it provides a standardized variant output as recommended by the Human Genome Variation Society.\n\nAvailabilityThe program is freely available online at http://bat.infspire.org/genomepd/glass/.\n\nContactnikos.darzentas@gmail.com, malcikova.jitka@fnbrno.cz\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Chromosome assembly of large and complex genomes using multiple references

Despite the rapid development of sequencing technologies, assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout, a reference-assisted assembly tool that now works for large and complex genomes. Taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. Using Ragout, we transformed NGS assemblies of 15 different Mus musculus and one Mus spretus genomes into sets of complete chromosomes, leaving less than 5% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long PacBio reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. Additionally, we applied Ragout to Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared to other genomes from the Muridae family. Chromosome color maps confirmed most large-scale rearrangements that Ragout detected.

bioinformatics

DNA sequence+shape kernel enables alignment-free modeling of transcription factor binding

MotivationTranscription factors (TFs) bind to specific DNA sequence motifs. Several lines of evidence suggest that TF-DNA binding is mediated in part by properties of the local DNA shape: the width of the minor groove, the relative orientations of adjacent base pairs, etc. Several methods have been developed to jointly account for DNA sequence and shape properties in predicting TF binding affinity. However, a limitation of these methods is that they typically require a training set of aligned TF binding sites.\n\nResultsWe describe a sequence+shape kernel that leverages DNA sequence and shape information to better understand protein-DNA binding preference and affinity. This kernel extends an existing class of k-mer based sequence kernels, based on the recently described di-mismatch kernel. Using three in vitro benchmark datasets, derived from universal protein binding microarrays (uPBMs), genomic context PBMs (gcPBMs) and SELEX-seq data, we demonstrate that incorporating DNA shape information improves our ability to predict protein-DNA binding affinity. In particular, we observe that (1) the k-spectrum+shape model performs better than the classical k-spectrum kernel, particularly for small k values; (2) the di-mismatch kernel performs better than the k-mer kernel, for larger k; and (3) the di-mismatch+shape kernel performs better than the di-mismatch kernel for intermediate k values.\n\nAvailabilityThe software is available at https://bitbucket.org/wenxiu/sequence-shape.git\n\nContactrohs@usc.edu, william-noble@uw.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Tumor Origin Detection with Tissue-Specific miRNA and DNA methylation Markers

MotivationCancer of unknown primary origin constitutes 3-5% of all human malignancies. Patients with these carcinomas present with metastases without an established primary site, which may not be found even by thorough histological search methods. Patients with cancer of unknown primary origin always have poor prognosis and hardly have efficient treatment since most cancers respond well to specific chemotherapy or hormone drugs. Many studies have proposed classifiers based on miRNAs or mRNAs to predict the tumor origins, but few study focus on high-dimensional DNA methylation profiles.\n\nResultsWe introduced three classifiers with novel feature selection algorithm combined with random forest to effectively identify highly tissue-specific epigenetics biomarkers such as microRNAs and CpG sites, which can help us predict the origin site of tumors. This algorithm, incorporating differential analysis and descending dimension algorithm, was applied on 14 histological tissues and over 5000 samples based on miRNA expression and DNA methylation profiles to assign given primary tumor to its origin tissue. Our study shows all of these three classifiers have an overall accuracy of 87.78% (72.55%-97.54%) based on miRNA datasets and an accuracy of 96.43% (MRMD: 87.85%-99.76%) or 97.06% (PCA: 92.44%-100%) based on DNA methylation datasets on predicting the origin of tumors and suggests that the biomarkers we selected can efficiently predict the origin of tumors and allow the clinicians to avoid adjuvant systemic therapy or to choose less aggressive therapeutic options. We also developed a user-friendly webserver which enables users to predict the origin site of tumors by uploading the miRNAs expression or DNA methylation profiles of those cancers.\n\nAvailabilityThe webserver, data, and code are accessible free of charge at http://server.malab.cn/MMCOP/\n\nContactzouquan@nclab.net\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

BioMake: a GNU Make-compatible utility for declarative workflow management

The Unix \"make\" program is widely used in bioinformatics pipelines, but suffers from problems that limit its application to large analysis datasets. These include reliance on file modification times to determine whether a target is stale, lack of support for parallel execution on clusters, and restricted flexibility to extend the underlying logic program. We present BioMake, a make-like utility that is compatible with most features of GNU Make and adds support for popular cluster-based job-queue engines, MD5 signatures as an alternative to timestamps, and logic programming extensions in Prolog. BioMake is available from https://github.com/evoldoers/biomake under the Creative Commons Attribution 3.0 US license. The only dependency is SWI-Prolog, available from http://www.swi-prolog.org/. Contact: Ian Holmes ihholmes+biomake@gmail.com or Chris Mungall cmungall+biomake@gmail.com.

bioinformatics