Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Using Optimal F-Measure and Random Resampling in Gene Ontology Enrichment Calculations

BackgroundA central question in bioinformatics is how to minimize arbitrariness and bias in analysis of patterns of enrichment in data. A prime example of such a question is enrichment of gene ontology (GO) classes in lists of genes. Our paper deals with two issues within this larger question. One is how to calculate the false discovery rate (FDR) within a set of apparently enriched ontologies, and the second how to set that FDR within the context of assessing significance for addressing biological questions, to answer these questions we compare a random resampling method with a commonly used method for assessing FDR, the Benjamini-Hochberg (BH) method. We further develop a heuristic method for evaluating Type II (false negative) errors to enable utilization of F-Measure binary classification theory for distinguishing \"significant\" from \"non-significant\" degrees of enrichment.\n\nResultsThe results show the preferability and feasibility of random resampling assessment of FDR over the analytical methods with which we compare it. They also show that the reasonableness of any arbitrary threshold depends strongly on the structure of the dataset being tested, suggesting that the less arbitrary method of F-measure optimization to determine significance threshold is preferable.\n\nConclusionTherefore, we suggest using F-measure optimization instead of placing an arbitrary threshold to evaluate the significance of Gene Ontology Enrichment results, and using resampling to replace analytical methods

bioinformatics

AQMM: Enabling Absolute Quantification of Metagenome and Metatranscriptome

Metatranscriptome has become increasingly important along with the application of next generation sequencing in the studies of microbial functional gene activity in environmental samples. However, the quantification of target active gene is hindered by the current relative quantification methods, especially when tracking the sharp environmental change. Great needs are here for an easy-to-perform method to obtain the absolute quantification. By borrowing information from the parallel metagenome, an absolute quantification method for both metagenomic and metatranscriptomic data to per gene/cell/volume/gram level was developed. The effectiveness of AQMM was validated by simulated experiments and was demonstrated with a real experimental design of comparing activated sludge with and without foaming. Our method provides a novel bioinformatic approach to fast and accurately conduct absolute quantification of metagenome and metatranscriptome in environmental samples. The AQMM can be accessed from https://github.com/biofuture/aqmm.

bioinformatics

Clinker: visualising fusion genes detected in RNA-seq data

Genomic profiling efforts have revealed a rich diversity of oncogenic fusion genes, and many are emerging as important therapeutic targets. While there are many ways to identify fusion genes from RNA-seq data, visualising these transcripts and their supporting reads remains challenging. Clinker is a bioinformatics tool written in Python, R and Bpipe, that leverages the superTranscript method to visualise fusion genes. We demonstrate the use of Clinker to obtain interpretable visualisations of the RNA-seq data that lead to fusion calls. In addition, we use Clinker to explore multiple fusion transcripts with novel breakpoints within the P2RY8-CRLF2 fusion gene in B-cell Acute Lymphoblastic Leukaemia (B-ALL).\n\nAvailability and ImplementationClinker is freely available from Github https://github.com/Oshlack/Clinker under a MIT License.\n\nContactalicia.oshlack@mcri.edu.au

bioinformatics

ChIPdig: a comprehensive user-friendly tool for mining multi-sample ChIP-seq data

BackgroundIn recent years, epigenetic research has enjoyed explosive growth as high-throughput sequencing technologies become more accessible and affordable. However, this advancement has not been matched with similar progress in data analysis capabilities from the perspective of experimental biologists not versed in bioinformatic languages. For instance, chromatin immunoprecipitation followed by next-generation sequencing (ChIP-seq) is at present widely used to identify genomic loci of transcription factor binding and histone modifications. Basic ChIP-seq data analysis, including read mapping and peak calling, can be accomplished through several well-established tools, but more sophisticated analyzes aimed at comparing data derived from different conditions or experimental designs constitute a significant bottleneck. We reason that the implementation of a single comprehensive ChIP-seq analysis pipeline could be beneficial for many experimental (wet lab) researchers who would like to generate genomic data.\n\nResultsHere we present ChIPdig, a stand-alone application with adjustable parameters designed to allow researchers to perform several analyzes, namely read mapping to a reference genome, peak calling, annotation of regions based on reference coordinates (e.g. transcription start and termination sites, exons, introns, 5' UTRs and 3' UTRs), and generation of heatmaps and metaplots for visualizing coverage. Importantly, ChIPdig accepts multiple ChIP-seq datasets as input, allowing genome-wide differential enrichment analysis in regions of interest to be performed. ChIPdig is written in R and enables access to several existing and highly utilized packages through a simple user interface powered by the Shiny package. Here, we illustrate the utility and user-friendly features of ChIPdig by analyzing H3K36me3 and H3K4me3 ChIP-seq profiles generated by the modENCODE project as an example.\n\nConclusionsChIPdig offers a comprehensive and user-friendly pipeline for analysis of multiple sets of ChIP-seq data by both experimental and computational researchers. It is open source and available at https://github.com/rmesse/ChIPdig.

bioinformatics

glactools: a command-line toolset for the management of genotype likelihoods and allele counts

MotivationResearch projects involving population genomics routinely need to store genotyping information, population allele frequencies, combine files from different samples, query the data and export it to various formats. This is often done using bespoke in-house scripts which cannot be easily adapted to new projects and seldom constitute reproducible workflows.\n\nResultsWe introduce glactools, a set of command-line utilities which can import data from genotypes or population-wide allele frequencies into an intermediate representation, compute various operations on it and export the data to several file formats used by population genetics software. This intermediate format can take 2 forms, one to store per-individual genotype likelihoods and a second for allele counts from one or more individuals. glactools allows users to perform operations such as intersecting datasets, merging individuals into populations, creating subsets, perform queries (e.g. return sites where a given population does not share an allele with a second one) and compute summary statistics to answer biologically relevant questions.\n\nAvailabilityglactools is freely available for use under the GPL. It requires a C++ compiler and the htslib library. (https://grenaud.github.io/glactools/).\n\nContactgabriel.reno@gmail.com\n\nSupplementary informationSupplementary methods and results are available at Bioinformatics online.

bioinformatics

High-throughput ANI Analysis of 90K Prokaryotic Genomes Reveals Clear Species Boundaries

A fundamental question in microbiology is whether there is a continuum of genetic diversity among genomes or clear species boundaries prevail instead. Answering this question requires robust measurement of whole-genome relatedness among thousands of genomes and from diverge phylogenetic lineages. Whole-genome similarity metrics such as Average Nucleotide Identity (ANI) can provide the resolution needed for this task, overcoming several limitations of traditional techniques used for the same purposes. Although the number of genomes currently available may be adequate, the associated bioinformatics tools for analysis are lagging behind these developments and cannot scale to large datasets. Here, we present a new method, FastANI, to compute ANI using alignment-free approximate sequence mapping. Our analyses demonstrate that FastANI produces an accurate ANI estimate and is up to three orders of magnitude faster when compared to an alignment (e.g., BLAST)-based approach. We leverage FastANI to compute pairwise ANI values among all prokaryotic genomes available in the NCBI database. Our results reveal a clear genetic discontinuity among the database genomes, with 99.8% of the total 8 billion genome pairs analyzed showing either >95% intra-species ANI or <83% inter-species ANI values. We further show that this discontinuity is recovered with or without the most frequently represented species in the database and is robust to historic additions in the public genome databases. Therefore, 95% ANI represents an accurate threshold for demarcating almost all currently named prokaryotic species, and wide species boundaries may exist for prokaryotes.

bioinformatics

SNPTB: nucleotide variant identification and annotation in Mycobacterium tuberculosis genomes

SummaryWhole genome sequencing (WGS) has become a mainstay in biomedical research. The continually decreasing cost of sequencing has resulted in a data deluge that underlines the need for easy-to-use bioinformatics pipelines that can mine meaningful information from WGS data. SNPTB is one such pipeline that analyzes WGS data originating from in vitro or clinical samples of Mycobacterium tuberculosis and outputs high-confidence single nucleotide polymorphisms in the bacterial genome. The name of the mutated gene and the functional consequence of the mutation on the gene product is also determined. SNPTB utilizes open source software for WGS data analyses and is written primarily for biologists with minimal computational skills.\n\nAvailability and implementationSNPTB is a python package and is available from https://github.com/aditi9783/SNPTB\n\nContactag1349@njms.rutgers.edu\n\nSupplementary informationTutorial for SNPTB is available at https://github.com/aditi9783/SNPTB/blob/master/docs/SNPTB_tutorial.md

bioinformatics

Barley Long Non-Coding RNAs and Their Tissue-Specific Co-expression Pattern with Coding-Transcripts

Long non-coding RNAs (lncRNA) with non-protein or small peptide-coding potential transcripts are emerging regulatory molecules. With the advent of next-generation sequencing technologies and novel bioinformatics tools, a tremendous number of lncRNAs has been identified in several plant species. Recent reports demonstrated roles of plant lncRNAs such as development and environmental response. Here, we reported a genome-wide discovery of ~8,000 barley lncRNAs and measured their expression pattern upon excessive boron (B) treatment. According to the tissue-based comparison, leaves have a greater number of B-responsive differentially expressed lncRNAs than the root. Functional annotation of the coding transcripts, which were co-expressed with lncRNAs, revealed that molecular function of the ion transport, establishment of localization, and response to stimulus significantly enriched only in the leaf. On the other hand, 32 barley endogenous target mimics (eTM) as lncRNAs, which potentially decoy the transcriptional suppression activity of 18 miRNAs, were obtained. Presented data including identification, expression measurement, and functional characterization of barley lncRNAs suggest that B-stress response might also be regulated by lncRNA expression via cooperative interaction of miRNA-eTM-coding target transcript modules.

bioinformatics

Comparative systems analysis of the secretome of the opportunistic pathogen Aspergillus fumigatus and other Aspergillus species

Aspergillus fumigatus and multiple other Aspergillus species cause a wide range of lung infections, collectively termed aspergillosis. Aspergilli are ubiquitous in environment with healthy immune systems routinely eliminating inhaled conidia, however, Aspergilli can become an opportunistic pathogen in immune-compromised patients. The aspergillosis mortality rate and emergence of drug-resistance reveals an urgent need to identify novel targets. Secreted and cell membrane proteins play a critical role in fungal-host interactions and pathogenesis. Using a computational pipeline integrating data from high-throughput experiments and bioinformatic predictions, we have identified secreted and cell membrane proteins in ten Aspergillus species known to cause aspergillosis. Small secreted and effector-like proteins similar to agents of fungal-plant pathogenesis were also identified within each secretome. A comparison with humans revealed that at least 70% of Aspergillus secretomes has no sequence similarity with the human proteome. An analysis of antigenic qualities of Aspergillus proteins revealed that the secretome is significantly more antigenic than cell membrane proteins or the complete proteome. Finally, overlaying an expression dataset, four A. fumigatus proteins upregulated during infection and with available structures, were found to be structurally similar to known drug target proteins in other organisms, and were able to dock in silico with the respective drug.

bioinformatics

Pathway enrichment analysis of -omics data

Pathway enrichment analysis helps gain mechanistic insight into large gene lists typically resulting from genome scale (-omics) experiments. It identifies biological pathways that are enriched in the gene list more than expected by chance. We explain pathway enrichment analysis and present a practical step-by-step guide to help interpret gene lists resulting from RNA-seq and genome sequencing experiments. The protocol comprises three major steps: define a gene list from genome scale data, determine statistically enriched pathways, and visualize and interpret the results. We focus on differentially expressed genes and mutated cancer genes, however the described principles can be applied to diverse -omics data. The protocol is designed for biologists with no prior bioinformatics training and uses freely available software including g:Profiler, GSEA, Cytoscape and Enrichment Map.

bioinformatics

MAGpy: a reproducible pipeline for the downstream analysis of metagenome-assembled genomes (MAGs)

Recent advances in bioinformatics have enabled the rapid assembly of genomes from metagenomes (MAGs), and there is a need for reproducible pipelines that can annotate and characterise thousands of genomes simultaneously. Here we present MAGpy, a Snakemake pipeline that takes FASTA input and compares MAGs to several public databases, checks quality, assigns a taxonomy and draws a phylogenetic tree.

bioinformatics

NanoPack: visualizing and processing long read sequencing data

Summary: Here we describe NanoPack, a set of tools developed for visualization and processing of long read sequencing data from Oxford Nanopore Technologies and Pacific Biosciences.\n\nAvailability and Implementation: The NanoPack tools are written in Python3 and released under the GNU GPL3.0 Licence. The source code can be found at https://github.com/wdecoster/nanopack, together with links to separate scripts and their documentation. The scripts are compatible with Linux, Mac OS and the MS Windows 10 subsystem for linux and are available as a graphical user interface, a web service at http://nanoplot.bioinf.be and command line tools.\n\nContact: wouter.decoster@molgen.vib-ua.be\n\nSupplementary information: Supplementary tables and figures are available at Bioinformatics online.

bioinformatics

Hybrid correction of highly noisy Oxford Nanoporelong reads using a variable-order de Bruijn graph

MotivationThe recent rise of long read sequencing technologies such as Pacific Biosciences and Oxford Nanopore allows to solve assembly problems for larger and more complex genomes than what allowed short reads technologies. However, these long reads are very noisy, reaching an error rate of around 10 to 15% for Pacific Biosciences, and up to 30% for Oxford Nanopore. The error correction problem has been tackled by either self-correcting the long reads, or using complementary short reads in a hybrid approach, but most methods only focus on Pacific Biosciences data, and do not apply to Oxford Nanopore reads. Moreover, even though recent chemistries from Oxford Nanopore promise to lower the error rate below 15%, it is still higher in practice, and correcting such noisy long reads remains an issue.\n\nResultsWe present HG-CoLoR, a hybrid error correction method that focuses on a seed-and-extend approach based on the alignment of the short reads to the long reads, followed by the traversal of a variable-order de Bruijn graph, built from the short reads. Our experiments show that HG-CoLoR manages to efficiently correct Oxford Nanopore long reads that display an error rate as high as 44%. When compared to other state-of-the-art long read error correction methods able to deal with Oxford Nanopore data, our experiments also show that HG-CoLoR provides the best trade-off between runtime and quality of the results, and is the only method able to efficiently scale to eukaryotic genomes.\n\nAvailability and implementationHG-CoLoR is implemented is C++, supported on Linux platforms and freely available at https://github.com/morispi/HG-CoLoR\n\nContact: pierre.morisse2@univ-rouen.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

SeqsLab: an integrated platform for cohort-based annotation and interpretation of genetic variants on Spark

SummarySeqsLab is a platform that helps researchers to easily annotate and interpret genetic variants derived from a large quantity of personal genomes. It provides an integrated interface to annotate the variants based on curated databases as well as in silico estimation on the effects of the variants. SeqsLab adopts the scalable cluster computing framework, Spark, and incorporates several customized algorithms to speed up the process of variant annotation and interpretation. The key features of SeqsLab include efficient annotation on large structural variations, diverse combinations of variant filters, easy incorporation with a vast amount of public databases, and scalable architecture of analyzing hundreds of human whole genomes simultaneously.\n\nAvailability and ImplementationSeqsLab is implemented with JAVA. The generated annotation will then be stored in Elasticsearch for real-time query and exploratory analysis. SeqsLab can be accessed by web browsers and is freely available at http://portal.seqslab.net/.\n\nContactchungtsai_su@atgenomix.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

MSAC: Compression of multiple sequence alignment files

MotivationBioinformatics databases grow rapidly and achieve values hardly to imagine a decade ago. Among numerous bioinformatics processes generating hundreds of GB is multiple sequence alignments of protein families. Its largest database, i.e., Pfam, consumes 40-230 GB, depending of the variant. Storage and transfer of such massive data has become a challenge.\n\nResultsWe propose a novel compression algorithm, MSAC (Multiple Sequence Alignment Compressor), designed especially for aligned data. It is based on a generalisation of the positional Burrows-Wheeler transform for non-binary alphabets. MSAC handles FASTA, as well as Stockholm files. It offers up to six times better compression ratio than other commonly used compressors, i.e., gzip. Performed experiments resulted in an analysis of the influence of a protein family size on the compression ratio.\n\nAvailabilityMSAC is available for free at https://github.com/refresh-bio/msac and http://sun.aei.polsl.pl/REFRESH/msac.\n\nContactsebastian.deorowicz@polsl.pl\n\nSupplementary materialSupplementary data are available at the publisher Web site.

bioinformatics

Characterization of strawberry (Fragaria vesca) sequence genome

BackgroundIn order to understand strawberry genes structure and evolution in this era of genomics, it is important to know the general statistical characteristics of the gene, intron and exon structures of strawberry and the expression of genes on different parts of strawberry genome. In the present study, about 32,422 genes on strawberry chromosomes were evaluated, and a number of bioinformatic softwares were used to analyze the characteristics of genes, exons and introns, expression of genes in different regions on the chromosomes. Also, the positions of strawberry centromeres were predicted.\n\nResultsOur results showed that, there are differences in the various features of different chromosomes and also vary in different parts of the same chromosome. The longer the number of genes, the longer the length of chromosome. The average length of genes is about 2809bp and the length of the individual gene is 0-2000bp with 5.3 exons and 4.3 introns per gene. The average length of the exon was 229bp and the intron was 413bp. Among the evaluated genes, ehe intronless gene accounted for 20.05%. Consistently a same trend with the expression levels of the same parts of the gene on a chromosome in different organizations was observed. Finally, the number of genes was positively correlated with the number of intronless, and there was a negative correlation of the length of the gene. The length of the gene depends primarily on the length of the intron, and the length of the exon has little effect on it. The number of exons was negatively correlated with the length of the exons, and the intron was also true.\n\nConclusionThe results of this investigation could definitely provide a significant foundation for further research on function analysis of gene family in Strawberry.

bioinformatics

Database-integrated genome screening (DIGS): exploring genomes heuristically using sequence similarity search tools and a relational database.

A significant fraction of most genomes is comprised of DNA sequences that have been incompletely investigated. This genomic dark matter contains a wealth of useful biological information that can be recovered by systematically screening genomes in silico using sequence similarity search tools. Specialized computational tools are required to implement these screens efficiently. Here, we describe the database-integrated genome-screening (DIGS) tool: a computational framework for performing these investigations. To demonstrate, we screen mammalian genomes for endogenous viral elements (EVEs) derived from the Filoviridae, Parvoviridae, Circoviridae and Bornaviridae families, identifying numerous novel elements in addition to those that have been described previously. The DIGS tool provides a simple, robust framework for implementing a broad range of heuristic, sequence analysis-based explorations of genomic diversity.\n\nAvailabilityhttp://giffordlabcvr.github.io/DIGS-tool/\n\nContactrobert.gifford@glasgow.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

miRge 2.0: An updated tool to comprehensively analyze microRNA sequencing data

miRNAs play important roles in the regulation of gene expression. The rapidly developing field of microRNA sequencing (miRNA-seq; small RNA-seq) needs comprehensive bioinformatics tools to analyze these large datasets. We present the second iteration of miRge, miRge 2.0, with multiple enhancements. miRge 2.0 adds new functionality including novel miRNA detection, A-to-I editing analysis, better output files, and improved alignment to miRNAs. Our novel miRNA detection method is the first to use both miRNA hairpin sequence structure and composition of isomiRs resulting in a more specific capture of potential miRNAs. Using known miRNA data, our support vector machine (SVM) model predicted miRNAs with an average Matthews correlation coefficient (MCC) of 0.939 over 32 human cell datasets and outperformed miRDeep2 and miRAnalyzer regarding phylogenetic conservation. The A-to-I editing analysis implementation strongly correlated with a reference datasets prior analysis with adjusted R2 = 0.96. miRge 2.0 comes with alignment libraries to both miRBase v21 and MirGeneDB for 6 species: human, mouse, rat, fruit fly, nematode and zebrafish; and has a tool to create custom libraries. With the redevelopment of the tool in Python, it is now incorporated into bcbio-nextgen and implementable through Bioconda. miRge 2.0 is freely available at: https://github.com/mhalushka/miRge.

bioinformatics