Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

SPECTRE: a Suite of PhylogEnetiC Tools for Reticulate Evolution

SummarySplit-networks are a generalization of phylogenetic trees that have proven to be a powerful tool in phylogenetics. Various ways have been developed for computing such networks, including split-decomposition, NeighborNet, QNet and FlatNJ. Some of these approaches are implemented in the user-friendly SplitsTree software package. However, to give the user the option to adjust and extend these approaches and to facilitate their integration into analysis pipelines, there is a need for robust, open-source implementations of associated data structures and algorithms. Here we present SPECTRE, a readily available, open-source library of data structures written in Java, that comes complete with new implementations of several pre-published algorithms and a basic interactive graphical interface for visualizing planar split networks. SPECTRE also supports the use of longer running algorithms by providing command line interfaces, which can be executed on servers or in High Performance Computing (HPC) environments.\n\nAvailabilityFull source code is available under the GPLv3 license at: https://github.com/maplesond/SPECTRE\n\nSPECTREs core library is available from Maven Central at: https://mvnrepository.com/artifactuk.ac.uea.cmp.spectre/core\n\nDocumentation is available at: http://spectre-suite-of-phylogenetic-tools-for-reticulate-evolution.readthedocs.io/en/latest/\n\nContactsarah.bastkowski@earlham.ac.uk\n\nSupplementary Information (SI)Supplementary information is available at Bioinformatics online.

bioinformatics

QuimP - Analyzing transmembrane signalling in highly deformable cells

SummaryTransmembrane signalling plays important physiological roles, with G protein-coupled cell surface receptors being particularly important therapeutic targets. Fluorescent proteins are widely used to study signalling, but the analysis of image time series can be challenging, in particular when changes in cell shape are involved. To this end we have developed QuimP software. QuimP semi-automatically tracks cell outlines, quantifies spatio-temporal patterns of fluorescence at the cell membrane, and tracks local shape deformations. QuimP is particularly useful for studying cell motility, for example in immune or cancer cells.\n\nAvailability and ImplementationQuimP (http://warwick.ac.uk/quimp) consists of a set of Java plugins for Fiji/ImageJ (http://fiji.sc/) and can be easily installed through the Fiji Updater (http://warwick.ac.uk/quimp/wiki-pages/installation). It is compatible with Mac, Windows and Unix-based operating systems, requiring version >1.45 of Fiji/ImageJ and Java 8. QuimP is released as open source (https://github.com/CellDynamics/QuimP/) under an academic licence.\n\nContactT.Bretschneider@warwick.ac.uk\n\nSupplementary InformationSupplementary materials (SI-A to SI-D) are available at Bioinformatics online. Test data is available from http://warwick.ac.uk/quimp/test_data.

bioinformatics

Integrative DNA copy number detection and genotyping from sequencing and array-based platforms

MotivationCopy number variations (CNVs) are gains and losses of DNA segments and have been associated with disease. Many large-scale genetic association studies are performing CNV analysis using whole exome sequencing (WES) and whole genome sequencing (WGS). In many of these studies, previous SNP-array data are available. An integrated cross-platform analysis is expected to improve resolution and accuracy, yet there is no tool for effectively combining data from sequencing and array platforms. The detection of CNVs using sequencing data alone can also be further improved by the utilization of allele-specific reads.\n\nResultsWe propose a statistical framework, integrated Copy Number Variation detection algorithm (iCNV), which can be applied to multiple study designs: WES only, WGS only, SNP array only, or any combination of SNP and sequencing data. iCNV applies platform specific normalization, utilizes allele specific reads from sequencing and integrates matched NGS and SNP-array data by a Hidden Markov Model (HMM). We compare integrated two-platform CNV detection using iCNV to naive intersection or union of platforms and show that iCNV increases sensitivity and robustness. We also assess the accuracy of iCNV on WGS data only, and show that the utilization of allele-specific reads improve CNV detection accuracy compared to existing methods.\n\nAvailabilityhttps://github.com/zhouzilu/iCNV\n\nContactnzh@wharton.upenn.edu, zhouzilu@mail.med.upenn.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

NeatSeq-Flow: A Lightweight Software for Efficient Execution of High Throughput Sequencing Workflows

Nowadays, it has become almost a necessity for many biologists to execute bioinformatics workflows (WFs) as part of their research. However, most WF-management software packages require for their operation at least some programming expertise. Here we describe NeatSeq-Flow, a platform that enables users with no programming knowledge to design and execute complex high throughput sequencing WFs. This is achieved by using a compendium of pre-built modules as well as a generic module, both do not require programming expertise. Nonetheless, NeatSeq-Flow retains the flexibility to generate sophisticated WF modules using templates and only basic Python programming abilities. NeatSeq-Flow is designed to enable easy sharing of WFs and modules by conceptually separating modules, WF design, sample information and execution. Moreover, NeatSeq-Flow works hand in hand with CONDA environments for easy installation of the WFs analysis programs in one go. NeatSeq-Flow enables efficient WF execution on computer clusters by parallelizing on both samples and WF steps. NeatSeq-Flow operates by shell-script generation; thus it allows full transparency of the WF process. NeatSeq-Flow offers real-time WF execution monitoring, detailed documentation and self-sustaining WF backups for reproducibility. All of these features make NeatSeq-Flow an easy-to-use WF platform while not compromising for flexibility, reproducibility, transparency and efficiency.\n\nAvailabilityhttp://neatseq-flow.readthedocs.io/en/latest/\n\nContactsklarz@bgu.ac.il

bioinformatics

COGEM: A Toolbox for Computational Genomics in Matlab

MotivationThe Matlab programming language is widely used for both teaching and research in engineering, computer science, and mathematics. Despite its many strengths, it has never been a dominant language in computational genomics or bioinformatics more generally.\n\nResultsHere, we introduce COGEM, a long-term project to develop computational genomics functionality in Matlab. The initial release provides functions for manipulating genomic intervals, stranded or unstranded, with or without numerical data associated. It includes features for both text and binary file input and output, conversion between BAM, BED and BEDGRAPH formats, and numerous functions for manipulating intervals, including shifting, expanding, overlapping, intersecting, unioning, finding nearest intervals, piling up intervals, and performing unary and binary numerical and logical operations on sets of intervals. The toolbox is well-suited to the analysis of high-throughput sequencing data. We demonstrate its functionality by creating a ChIP-seq peak-calling algorithm by chaining together a series of commands, and find it capable of analyzing genome-scale data in reasonable time.\n\nAvailabilityThe current toolbox and reference manual is available as supplementary material, and updated versions will be maintained at www.perkinslab.ca online.

bioinformatics

CarrierSeq: a sequence analysis workflow for low-input nanopore sequencing

MotivationLong-read nanopore sequencing technology is of particular significance for taxonomic identification at or below the species level. For many environmental samples, the total extractable DNA is far below the current input requirements of nanopore sequencing, preventing \"sample to sequence\" metagenomics from low-biomass or recalcitrant samples.\n\nResultsHere we address this problem by employing carrier sequencing, a method to sequence low-input DNA by preparing the target DNA with a genomic carrier to achieve ideal library preparation and sequencing stoichiometry without amplification. We then use CarrierSeq, a sequence analysis workflow to identify the low-input target reads from the genomic carrier. We tested CarrierSeq experimentally by sequencing from a combination of 0.2 ng Bacillus subtilis ATCC 6633 DNA in a background of 1 g Enterobacteria phage {lambda} DNA. After filtering of carrier, low quality, and low complexity reads, we detected target reads (B. subtilis), contamination reads, and \"high quality noise reads\" (HQNRs) not mapping to the carrier, target or known lab contaminants. These reads appear to be artifacts of the nanopore sequencing process as they are associated with specific channels (pores). By treating reads as a Poisson arrival process, we implement a statistical test to reject data from channels dominated by HQNRs while retaining target reads.\n\nAvailabilityCarrierSeq is an open-source bash script with supporting python scripts which leverage a variety of bioinformatics software packages on macOS and Ubuntu. Supplemental documentation is available from Github - https://github.com/amojarro/carrierseq. In addition, we have compiled all required dependencies in a Docker image available from - https://hub.docker.com/r/mojarro/carrierseq.

bioinformatics

LOGAN: A framework for LOssless Graph-based ANalysis of high throughput sequence data

Recent massive growth in the production of sequencing data necessitates matching improve-ments in bioinformatics tools to effectively utilize it. Existing tools suffer from limitations in both scalability and applicability which are inherent to their underlying algorithms and data structures. We identify the key requirements for the ideal data structure for sequence analy-ses: it should be informationally lossless, locally updatable, and memory efficient; requirements which are not met by data structures underlying the major assembly strategies Overlap Layout Consensus and De Bruijn Graphs. We therefore propose a new data structure, the LOGAN graph, which is based on a memory efficient Sparse De Bruijn Graph with routing information. Innovations in storing routing information and careful implementation allow sequence datasets for Escherichia coli (4.6Mbp, 117x coverage), Arabidopsis thaliana (135Mbp, 17.5x coverage) and Solanum pennellii (1.2Gbp, 47x coverage) to be loaded into memory on a desktop computer in seconds, minutes, and hours respectively. Memory consumption is competitive with state of the art alternatives, while losslessly representing the reads in an indexed and updatable form. Both Second and Third Generation Sequencing reads are supported. Thus, the LOGAN graph is positioned to be the backbone for major breakthroughs in sequence analysis such as integrated hybrid assembly, assembly of exceptionally large and repetitive genomes, as well as assembly and representation of pan-genomes.

bioinformatics

The Oyster River Protocol: A Multi Assembler and Kmer Approach For de novo Transcriptome Assembly

Characterizing transcriptomes in non-model organisms has resulted in a massive increase in our understanding of biological phenomena. This boon, largely made possible via high-throughput sequencing, means that studies of functional, evolutionary and population genomics are now being done by hundreds or even thousands of labs around the world. For many, these studies begin with a de novo transcriptome assembly, which is a technically complicated process involving several discrete steps. The Oyster River Protocol (ORP), described here, implements a standardized and benchmarked set of bioinformatic processes, resulting in an assembly with enhanced qualities over other standard assembly methods. Specifically, ORP produced assemblies have higher TransRate scores and mapping rates, which is largely a product of the fact that it leverages a multi-assembler and kmer assembly process, thereby bypassing the shortcomings of any one approach. These improvements are important, as previously unassembled transcripts are included in ORP assemblies, resulting in a significant enhancement of the power of downstream analysis. Further, as part of this study, we show that assembly quality is unrelated to taxonomy, nor is it related to the number of reads generated, above 30 million reads.\n\nCode AvailabilityThe version controlled open-source code is available at https://github.com/macmanes-lab/Oyster_River_Protocol. Instructions for software installation and use, and other details are available at http://oyster-river-protocol.rtfd.org/.

bioinformatics

GIC: A computational method for predicting the essentiality of long noncoding lncRNAs

Measuring the essentiality of genes is critically important in biology and medicine. Some bioinformatic methods have been developed for this issue but none of them can be applied to long noncoding RNAs (lncRNAs), one big class of biological molecules. Here we developed a computational method, GIC (Gene Importance Calculator), which can predict the essentiality of both protein-coding genes and lncRNAs based on RNA sequence information. For identifying the essentiality of protein-coding genes, GIC is competitive with well-established computational scores. More important, GIC showed a high performance for predicting the essentiality of lncRNAs. In an independent mouse lncRNA dataset, GIC achieved an exciting performance (AUC=0.918). In contrast, the traditional computational methods are not applicable to lncRNAs. As a public web server, GIC is freely available at http://www.cuilab.cn/gic/.

bioinformatics

CHIC: a short read aligner for pan-genomic references

Recently the topic of computational pan-genomics has gained increasing attention, and particularly the problem of moving from a single-reference paradigm to a pan-genomic one. Perhaps the simplest way to represent a pan-genome is to represent it as a set of sequences. While indexing highly repetitive collections has been intensively studied in the computer science community, the research has focused on efficient indexing and exact pattern patching, making most solutions not yet suitable to be used in bioinformatic analysis pipelines.\n\nResultsWe present CHIC, a short-read aligner that indexes very large and repetitive references using a hybrid technique that combines Lempel-Ziv compression with Burrows-Wheeler read aligners.\n\nAvailabilityOur tool is open source and available online at https://gitlab.com/dvalenzu/CHIC

bioinformatics

R2DGC: Threshold-free peak alignment and identification for 2D gas chromatography mass in R

Summary: Comprehensive two dimensional gas chromatography-mass spectrometry is a powerful method for analyzing complex mixtures of volatile compounds. This method produces a large amount of raw data that requires downstream processing to align signals of interest (peaks) across multiple samples and match peak characteristics to reference standard libraries prior to downstream statistical analysis. To address the paucity of applications addressing this need, we have developed an R package that implements retention time and mass spectra similarity threshold-free alignments, seamlessly integrates retention time standards for universally reproducible alignments, performs common ion filtering, and provides compatibility with multiple peak quantification methods. We demonstrate the packages utility on a controlled mix of metabolite standards separated under variable chromatography conditions and data generated from cell lines.\n\nAvailability and documentation: R2DGC can be downloaded at https://github.com/rramaker/R2DGC or installed via the Comprehensive R Archive Network (CRAN).\n\nContact: sjcooper@hudsonalpha.org\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics

Skip-mers: increasing entropy and sensitivity to detect conserved genic regions with simple cyclic q-grams

Bioinformatic analyses and tools make extensive use of k-mers (fixed contiguous strings of k nucleotides) as an informational unit. K-mer analyses are both useful and fast, but are strongly affected by single-nucleotide polymorphisms or sequencing errors, effectively hindering direct-analyses of whole regions and decreasing their usability between evolutionary distant samples.\n\nWe introduce a concept of skip-mers, a cyclic pattern of used-and-skipped positions of k nucleotides spanning a region of size S [≥] k, and show how analyses are improved compared to using k-mers. The entropy of skip-mers increases with the larger span, capturing information from more distant positions and increasing the specificity, and uniqueness, of larger span skip-mers within a genome. In addition, skip-mers constructed in cycles of 1 or 2 nucleotides in every 3 (or a multiple of 3) lead to increased sensitivity in the coding regions of genes, by grouping together the more conserved nucleotides of the protein-coding regions.\n\nWe implemented a set of tools to count and intersect skip-mers between different datasets. We used these tools to show how skip-mers have advantages over k-mers in terms of entropy and increased sensitivity to detect conserved coding sequence, allowing better identification of genic matches between evolutionarily distant species. We also highlight potential applications to problems such as whole-genome alignment and multi-genome evolutionary analyses.\n\nSoftware availabilitythe skm-tools implementing the methods described in this manuscript are available under MIT license at http://github.com/bioinfologics/skm-tools/

bioinformatics

B-cell receptor reconstruction from single-cell RNA-seq with VDJPuzzle

The B-cell receptor (BCR) performs essential functions for the adaptive immune system including recognition of pathogen-derived antigens. Cell-to-cell variability of BCR sequences due to V(D)J recombination and somatic hypermutation (SHM) necessitates single-cell characterization of BCR sequences. Single-cell RNA sequencing (scRNA-seq) presents the opportunity for simultaneous capture of the BCR sequence and transcriptomic signature for a detailed understanding of the dynamics of an immune response.\n\nWe developed VDJPuzzle 2.0, a bioinformatic tool that reconstructs productive, full-length B-cell receptor sequences of both heavy and light chains. VDJPuzzle successfully reconstructs BCRs from 98.3% (n=117) of human and 96.5% (n=200) from murine B cells. 92.0% of clonotypes and 90.3% of mutations were concordant with single-cell Sanger sequencing of the immunoglobulin chains. VDJPuzzle is available at https://bitbucket.org/kirbyvisp/vdjpuzzle2

bioinformatics

BITE: an R package for biodiversity analyses

Nowadays, molecular data analyses for biodiversity studies often require advanced bioinformatics skills, preventing many life scientists from analyzing their own data autonomously. BITE R package provides complete and user-friendly functions to handle SNP data and third-party software results (i.e. Admixture, TreeMix), facilitating their visualization, interpretation and use. Furthermore, BITE implements additional useful procedures, such as representative sampling and bootstrap for TreeMix, filling the gap in existing biodiversity data analysis tools.\n\nAvailabilityhttps://github.com/marcomilanesi/BITE

bioinformatics

GenomeDISCO: A concordance score for chromosome conformation capture experiments using random walks on contact map graphs

MotivationThe three-dimensional organization of chromatin plays a critical role in gene regulation and disease. High-throughput chromosome conformation capture experiments such as Hi-C are used to obtain genome-wide maps of 3D chromatin contacts. However, robust estimation of data quality and systematic comparison of these contact maps is challenging due to the multi-scale, hierarchical structure of chromatin contacts and the resulting properties of experimental noise in the data. Measuring concordance of contact maps is important for assessing reproducibility of replicate experiments and for modeling variation between different cellular contexts.\n\nResultsWe introduce a concordance measure called GenomeDISCO (DIfferences between Smoothed COntact maps) for assessing the similarity of a pair of contact maps obtained from chromosome conformation capture experiments. The key idea is to smooth contact maps using random walks on the contact map graph, before estimating concordance. We use simulated datasets to benchmark GenomeDISCOs sensitivity to different types of noise that affect chromatin contact maps. When applied to a large collection of Hi-C datasets, GenomeDISCO accurately distinguishes biological replicates from samples obtained from different cell types. GenomeDISCO also generalizes to other chromosome conformation capture assays, such as HiChIP.\n\nAvailabilitySoftware implementing GenomeDISCO is available at https://github.com/kundajelab/genomedisco.\n\nContactakundaje@stanford.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

accuMUlate: A mutation caller designed for mutation accumulation experiments

MotivationMutation accumulation (MA) is the most widely used method for directly studying the effects of mutation. Modern sequencing technologies have led to an increased interest in MA experiments. By sequencing whole genomes from MA lines, researchers can directly study the rate and molecular spectra of spontaneous mutations and use these results to understand how mutation contributes to biological processes. At present there is no software designed specifically for identifying mutations from MA lines. Studies that combine MA with whole genome sequencing use custom bioinformatic pipelines that implement heuristic rules to identify putative mutations.\n\nResultsHere we describe O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW, a program that is designed to detect mutations from MA experiments. O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW implements a probabilistic model that reflects the design of a typical MA experiments while being flexible enough to accommodate properties unique to any particular experiment. For each putative mutation identified from this model O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW calculates a set of summary statistics that can be used to filter sites that may be false positives. A companion tool, O_SCPLOWDENOMINATEC_SCPLOW, can be used to apply filtering rules based on these statistics to simulated mutations and thus identify the number of callable sites per sample.\n\nAvailabilitySource code and releases available from https://github.com/dwinter/accuMUlate.

bioinformatics

Recovering genomic clusters of secondary metabolites from lakes: a Metagenomics 2.0 approach

BackgroundMetagenomic approaches became increasingly popular in the past decades due to decreasing costs of DNA sequencing and bioinformatics development. So far, however, the recovery of long genes coding for secondary metabolism still represents a big challenge. Often, the quality of metagenome assemblies is poor, especially in environments with a high microbial diversity where sequence coverage is low and complexity of natural communities high. Recently, new and improved algorithms for binning environmental reads and contigs have been developed to overcome such limitations. Some of these algorithms use a similarity detection approach to classify the obtained reads into taxonomical units and to assemble draft genomes. This approach, however, is quite limited since it can classify exclusively sequences similar to those available (and well classified) in the databases.\n\nIn this work, we used draft genomes from Lake Stechlin, north-eastern Germany, recovered by MetaBat, an efficient binning tool that integrates empirical probabilistic distances of genome abundance, and tetranucleotide frequency for accurate metagenome binning. These genomes were screened for secondary metabolism genes, such as polyketide synthases (PKS) and non-ribosomal peptide synthases (NRPS), using the Anti-SMASH and NAPDOS workflows.\n\nResultsWith this approach we were able to identify 243 secondary metabolite clusters from 121 genomes recovered from the lake samples. A total of 18 NRPS, 19 PKS and 3 hybrid PKS/NRPS clusters were found. In addition, it was possible to predict the partial structure of several secondary metabolite clusters allowing for taxonomical classifications and phylogenetic inferences.\n\nConclusionsOur approach revealed a great potential to recover and study secondary metabolites genes from any aquatic ecosystem.

bioinformatics

Understanding and mitigating some limitations of Illumina (C) MiSeq for environmental sequencing of fungi.

ITS-amplicon metabarcode studies using the illumina MiSeq sequencing platform are the current standard tool for fungal ecology studies. Here we report on some of the particular challenges experienced while creating and using a ribosomal RNA gene (rDNA) amplicon library for an ecological study. Two significant complications were encountered. First, artificial differences in read abundances among OTUs were observed, apparently resulting from bias at two stages: PCR amplification of genomic DNA with ITS-region Illumina-sequence-adapted-primers, and during Illumina sequencing. These differential read abundances were only partially corrected by a common variance-stabilization method. Second, tag-switching (or the shifting of amplicons to incorrect sample indices) occurred at high levels in positive mock-community controls. An example of a bioinformatic method to estimate the rate of tag switching is shown, some recommendations on the use of positive controls and primer choice are given, and one approach to reducing potential false positives resulting from these technological biases is presented.

bioinformatics