Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Phylogeny-aware Identification and Correction of Taxonomically Mislabeled Sequences

Molecular sequences in public databases are mostly annotated by the submitting authors without further validation. This procedure can generate erroneous taxonomic sequence labels. Mislabeled sequences are hard to identify, and they can induce downstream errors because new sequences are typically annotated using existing ones. Furthermore, taxonomic mislabelings in reference sequence databases can bias metagenetic studies which rely on the taxonomy. Despite significant efforts to improve the quality of taxonomic annotations, the curation rate is low because of the labour-intensive manual curation process.\n\nHere, we present SATIVA, a phylogeny-aware method to automatically identify taxonomically mislabeled sequences ( mislabels) using statistical models of evolution. We use the Evolutionary Placement Algorithm (EPA) to detect and score sequences whose taxonomic annotation is not supported by the underlying phylogenetic signal, and automatically propose a corrected taxonomic classification for those. Using simulated data, we show that our method attains high accuracy for identification (96.9% sensitivity / 91.7% precision) as well as correction (94.9% sensitivity / 89.9% precision) of mislabels. Furthermore, an analysis of four widely used microbial 16S reference databases (Greengenes, LTP, RDP and SILVA) indicates that they currently contain between 0.2% and 2.5% mislabels. Finally, we use SATIVA to perform an in-depth evaluation of alternative taxonomies for Cyanobacteria.\n\nSATIVA is freely available at https://github.com/amkozlov/sativa.

Bioinformatics

The Ensembl Variant Effect Predictor

The Ensembl Variant Effect Predictor (VEP) is a powerful toolset for the analysis, annotation and prioritization of genomic variants, including in non-coding regions.\n\nThe VEP accurately predicts the effects of sequence variants on transcripts, protein products, regulatory regions and binding motifs by leveraging the high quality, broad scope, and integrated nature of the Ensembl databases. In addition, it enables comparison with a large collection of existing publicly available variation data within Ensembl to provide insights into population and ancestral genetics, phenotypes and disease.\n\nThe VEP is open source and free to use. It is available via a simple web interface (http://www.ensembl.org/vep), a powerful downloadable package, and both Ensembls Perl and REST application program interface (API) services.

Bioinformatics

Exploring miRNAs as the key to understand symptoms induced by ZIKA virus infection through a collaborative database.

Zika virus (ZIKV) is an emerging mosquito-borne flavivirus, first isolated in 1947 from the serum of a pyrexial rhesus monkey caged in the Zika Forest (Uganda/Africa)1. In 2007 ZIKV was reported to of been responsible for an outbreak of relatively mild disease, characterized by rash, arthralgia, and conjunctivitis on Yap Island, in the western Pacific Ocean2. In the past year, ZIKV has been circulating in the Americas, probably introduced through Easter Island (Chile), by French Polynesians3. In early 2015, a new outbreak was recognized in northeast Brazil4, where concerns over its possible links with infant microcephaly have been discussed5,6. Providing a definitive link between ZIKV infection and birth defects is still a big challenge7. MicroRNAs (miRNAs), are small noncoding RNAs that regulating post-transcriptional gene expression by translational repression, and play important roles in viral pathogenesis8 and brain development9. It is estimated that more than 60% of human protein-coding genes contain at least one conserved miRNA-binding site10. The potential for flavivirus-mediated miRNA signaling dysfunction in brain-tissue develop provides a compelling mechanism underlying perceived linked between ZIKV and microcephaly. Here, we provide strong evidences toward to understand the mechanism in which miRNAs can be linked to the \"congenital Zika syndrome\" symptoms. Moreover, following World Health Organization (WHO) recommendations11, we have assembled a database that could help target mechanistic investigations of this possible relationship between ZIKV symptoms and miRNA mediated human gene expression, helping to foster potential targets for therapy.

Bioinformatics

Combining multiple tools outperforms individual methods in gene set enrichment analyses

Gene set enrichment (GSE) analysis allows researchers to efficiently extract biological insight from long lists of differentially expressed genes by interrogating them at a systems level. In recent years, there has been a proliferation of GSE analysis methods and hence it has become increasingly difficult for researchers to select an optimal GSE tool based on their particular data set. Moreover, the majority of GSE analysis methods do not allow researchers to simultaneously compare gene set level results between multiple experimental conditions.\n\nResults: The ensemble of genes set enrichment analyses (EGSEA) is a method developed for RNA-sequencing data that combines results from twelve algorithms and calculates collective gene set scores to improve the biological relevance of the highest ranked gene sets. redEGSEAs gene set database contains around 25,000 gene sets from sixteen collections. It has multiple visualization capabilities that allow researchers to view gene sets at various levels of granularity. EGSEA has been tested on simulated data and on a number of human and mouse data sets and, based on biologists' feedback, consistently outperforms the individual tools that have been combined. Our evaluation demonstrates the superiority of the ensemble approach for GSE analysis, and its utility to effectively and efficiently extrapolate biological functions and potential involvement in disease processes from lists of differentially regulated genes.\n\nAvailability and Implementation: EGSEA is available as an R package at http://www.bioconductor.org/packages/EGSEA/. The gene sets collections are available in the R package EGSEAdata from http://www.bioconductor.org/packages/EGSEA/.

Bioinformatics

mLDM: a new hierarchical Bayesian statistical model for sparse microbioal association discovery

Interpretive analysis of metagenomic data depends on an understanding of the underlying associations among microbes from metagenomic samples. Although several statistical tools have been developed for metage-nomic association studies, they suffer from compositional bias or fail to take into account environmental factors that directly affect the composition of a given microbial community. In this paper, we propose metagenomic Lognormal-Dirichlet-Multinomial (mLDM), a hierarchical Bayesian model with sparsity constraints to bypass compositional bias and discover new associations among microbes and between microbes and environmental factors. The mLD-M model can 1) infer both conditionally dependent associations among microbes and direct associations between microbes and environmental factors; 2) consider both compositional bias and variance of metagenomic data; and 3) estimate absolute abundance for microbes. Thus, conditionally dependent association can capture direct relationship underlying microbial pairs and remove the indirect connections induced from other common factors. Empirical studies show the effectiveness of the mLDM model, using both synthetic data and the TARA Oceans eukaryotic data by comparing it with several state-of-the-art methodologies. Finally, mLDM is applied to western English Channel data and finds some interesting associations.

Bioinformatics

AnnoLnc: a web server for systematically annotating novel human lncRNAs

Although the repertoire of human lncRNAs has rapidly expanded, their biological function and regulation remain largely elusive. Here, we present AnnoLnc (http://annolnc.cbi.pku.edu.cn), an online portal for systematically annotating newly identified human lncRNAs. AnnoLnc offers a full spectrum of annotations covering genomic location, RNA secondary structure, expression, transcriptional regulation, miRNA interaction, protein interaction, genetic association and evolution, as well as an abstraction-based text summary and various intuitive figures to help biologists quickly grasp the essentials. In addition to an intuitive and mobile-friendly Web interactive design, AnnoLnc supports batch analysis and provides JSON-based Web Service APIs for programmatic analysis. To the best of our knowledge, AnnoLnc is the first web server to provide on-the-fly and systematic annotation for newly identified human lncRNAs. Some case studies have shown the power of AnnoLnc to inspire novel hypotheses.

Bioinformatics

Discovery of Primary, Cofactor, and Novel Transcription Factor Binding Site Motifs by Recursive, Thresholded Entropy Minimization

Data from ChIP-seq experiments can determine the genome-wide binding specificities of transcription factors (TFs) and other regulatory proteins. In the present study, we analyzed 745 ENCODE ChIP-seq peak datasets of 189 human TFs with a novel motif discovery method that is based on recursive, thresholded entropy minimization. This method is able to distinguish correct information models from noisy motifs, quantify the strengths of individual sites based on affinity, and detect adjacent cofactor binding sites that coordinate with primary TFs. We derived homogeneous and bipartite information models for 89 sequence-specific TFs, which enabled discovery of 24 cofactor motifs for 118 TFs, and revealed 6 high-confidence novel motifs. The reliability and accuracy of these models were determined via three independent quality control criteria, including the detection of experimentally proven binding sites, comparison with previously published motifs and statistical analyses. We also predict previously unreported TF cobinding interactions, and new components of known TF complexes. Because they are based on information theory, the derived models constitute a powerful tool for detecting and predicting the effects of variants in known binding sites, and predicting previously unrecognized binding sites and target genes.

Bioinformatics

A fast and accurate method for detection of IBD shared haplotypes in genome-wide SNP data

Identical by descent (IBD) segments are used to understand a number of fundamental issues in genetics. IBD segments are typically detected using long stretches of identical alleles between haplotypes in whole-genome SNP data. Phase or SNP call errors in genomic data can degrade accuracy of IBD detection and lead to false positive calls, false negative calls, and under- or overextension of true IBD segments. Furthermore, the number of comparisons increases quadratically with sample size, requiring high computational efficiency. We developed a new IBD segment detection program, FISHR (Find IBD Shared Haplotypes Rapidly), in an attempt to accurately detect IBD segments and to better estimate their endpoints using an algorithm that is fast enough to be deployed on the very large whole-genome SNP datasets. We compared the performance of FISHR to three leading IBD segment detection programs: GERMLINE, refinedIBD, and HaploScore. Using simulated and real genomic sequence data, we show that FISHR is slightly more accurate than all programs at detecting long (>3 cM) IBD segments but slightly less accurate than refinedIBD at detecting short (~1 cM) IBD segments. Moreover, FISHR outperforms all programs in determining the true endpoints of IBD segments, which is important for several reasons. FISHR takes two to four times longer than GERMLINE to run, whereas both GERMLINE and FISHR were orders of magnitude faster than refinedIBD and HaploScore. Overall, FISHR provides accurate IBD detection in unrelated individuals and is computationally efficient enough to be utilized on large SNP datasets > 20,000 individuals.

Bioinformatics

Crunch: Completely Automated Analysis of ChIP-seq Data

Although it has become routine for experimental groups to apply ChIP-seq technology to quantitatively characterize the genome-wide binding of transcription factors (TFs), computational analysis procedures remain far from standardized, making it difficult to meaningfully compare ChIP-seq results across experiments. In addition, while genome-wide binding patterns must ultimately be determined by local constellations of binding sites in the DNA, current analysis is typically limited to a standard search for enriched motifs in ChIP-seq peaks.\n\nHere we present Crunch, a completely automated computational method that performs all ChIP-seq analysis from quality control through read mapping and peak detecting, and integrates comprehensive modeling of the ChIP signal in terms of known and novel binding motifs, quantifying the contribution of each motif, and annotating which combinations of motifs explain each binding peak.\n\nApplying Crunch to 128 ChIP-seq datasets from the ENCODE project we find that TFs naturally separate into solitary TFs, for which a single motif explains the ChIP-peaks, and co-binding TFs for which multiple motifs co-occur within peaks. Moreover, for most datasets the motifs that Crunch identified de novo outperform known motifs and both the set of co-binding motifs and the top motif of solitary TFs are consistent across experiments and cell lines. Crunch is implemented as a web server (crunch.unibas.ch), enabling standardized analysis of any collection of ChIP-seq datasets by simply uploading raw sequencing data. Results are provided both in a graphical interface and as downloadable files.

Bioinformatics

Highly accessible AU-rich regions in 3′ untranslated regions are hotspots for binding of proteins and miRNAs

BackgroundMicroRNAs (miRNAs) are endogenous short non-coding RNAs involved in the regulation of gene expression at the post-transcriptional level typically by promoting destabilization or translational repression of target RNAs. Sometimes this regulation is absent or different, which likely is the result of interactions with other post-transcriptional factors, particularly RNA-binding proteins (RBPs). Despite the importance of the interactions between RBPs and miRNAs, little is known about how they affect post-transcriptional regulation in a global scale.\n\nResultsIn this study, we have analyzed CLIP datasets of 49 RBPs in HEK293 cells with the aim of understanding the interplay between RBPs and miRNAs in post-transcriptional regulation. Our results show that RBPs bind preferentially in conserved regulatory hotspots that frequently contain miRNA target sites. This organization facilitates the competition and cooperation among RBPs and the regulation of miRNA target site accessibility. In some cases RBP enrichment on target sites correlates with miRNA expression, suggesting coordination between the regulatory factors. However, in most cases, competition among factors is the most plausible interpretation of our data. Upon AGO2 knockdown, transcripts that contain such hotspots that overlap target sites of expressed miRNAs in 3UTRs are significantly less up-regulated than transcripts without them, suggesting that RBP binding limits miRNA accessibility.\n\nConclusionsWe show that RBP binding is concentrated in regulatory hotspots in 3UTRs. The presence of these hotspots facilitates the interaction among post-transcriptional regulators, that interact or compete with each other under different conditions. These hotspots are enriched in genes with regulatory functions such as DNA binding and RNA binding. Taken together, our results suggest that hotspots are important regulatory regions that define an extra layer of auto-regulatory control of post-transcriptional regulation.

Bioinformatics

A method for downstream analysis of gene set enrichment results facilitates the biological interpretation of vaccine efficacy studies

Gene set enrichment analysis (GSEA) is a widely employed method for analyzing gene expression profiles. The approach uses annotated sets of genes, identifies those that are coordinately up- or down-regulated in a biological comparison of interest, and thereby elucidates underlying biological processes relevant to the comparison. As the number of gene sets available in various collections for enrichment analysis has grown, the resulting lists of significant differentially regulated gene sets may also become larger, leading to the need for additional downstream analysis of GSEA results. Here we present a method that allows the rapid identification of a small number of co-regulated groups of genes - \"leading edge metagenes\" (LEMs) - from high scoring sets in GSEA results. LEM are sub-signatures which are common to multiple gene sets and that \"explain\" their enrichment specific to the experimental dataset of interest. We show that LEMs contain more refined lists of context-dependent and biologically meaningful genes than the parental gene sets. LEM analysis of the human vaccine response using a large database of immune signatures identified core biological processes induced by five different vaccines in datasets from human peripheral blood mononuclear cells (PBMC). Further study of these biological processes over time following vaccination showed that at day 3 post-vaccination, vaccines derived from viruses or viral subunits exhibit patterns of biological processes that are distinct from protein conjugate vaccines; however, by day 7 these differences were less pronounced. This suggests that the immune response to diverse vaccines eventually converge to a common transcriptional response. LEM analysis can significantly reduce the dimensionality of enriched gene sets, improve the identification of core biological processes active in a comparison of interest, and simplify the biological interpretation of GSEA results.\n\nAuthor SummaryGenome-wide expression profiling is a widely used tool to identify biological mechanisms in a comparison of interest. One analytic method, Gene set enrichment analysis (GSEA) uses annotated sets of genes and identifies those that are coordinately up- or down-regulated in a biological comparison of interest. This approach capitalizes on the fact that alternations in biological processes often cause the coordinated change of a large number of genes. However, as the number of gene sets available in various collections for enrichment analysis has grown, the resulting lists of significant differentially regulated gene sets may also become larger, leading to the need for additional downstream analysis of GSEA results. Here we present a method that allows the identification of a small number of co-regulated groups of genes - \"leading edge metagenes\" (LEMs) - from high scoring sets in GSEA results. We show that LEMs contain more refined lists of context-dependent biologically meaningful genes than the parental gene sets and demonstrate the utility of this approach in analyzing the transcriptional response to vaccination. LEM analysis can significantly reduce the dimensionality of enriched gene sets, improve the identification of core biological processes active in a comparison of interest, and facilitate the biological interpretation of GSEA results.

Bioinformatics

RSQ: a statistical method for quantification of isoform-specific structurome using transcriptome-wide structural profiling data

The structure of RNA, which is considered to be a second layer of information alongside the genetic code, provides fundamental insights into the cellular function of both coding and non-coding RNAs. Several high-throughput technologies have been developed to profile transcriptome-wide RNA structures, i.e., the structurome. However, it is challenging to interpret the profiling data because the observed data represent an average over different RNA conformations and isoforms with different abundance. To address this challenge, we developed an RNA structurome quantification method (RSQ) to statistically model the distribution of reads over both isoforms and RNA conformations, and thus provide accurate quantification of the isoform-specific structurome. The quantified RNA structurome enables the comparison of isoform-specific conformations between different conditions, the exploration of RNA conformation variation affected by single nucleotide polymorphism (SNP),and the measurement of RNA accessibility for binding of either small RNAs in RNAi-based assays or RNA binding protein in transcriptional regulation. The model used in our method sheds new light on the potential impact of the RNA structurome on gene regulation.

Bioinformatics

From genomes to phenotypes: Traitar, the microbial trait analyzer

The number of sequenced genomes is growing exponentially, profoundly shifting the bottleneck from data generation to genome interpretation. Traits are often used to characterize and distinguish bacteria, and are likely a driving factor in microbial community composition, yet little is known about the traits of most microbes. We describe Traitar, the microbial trait analyzer, which is a fully automated software package for deriving phenotypes from the genome sequence. Traitar provides phenotype classifiers to predict 67 traits related to the use of various substrates as carbon and energy sources, oxygen requirement, morphology, antibiotic susceptibility, proteolysis and enzymatic activities. Furthermore, it suggests protein families associated with the presence of particular phenotypes. Our method uses L1-regularized L2-loss support vector machines for phenotype assignments based on phyletic patterns of protein families and their evolutionary histories across a diverse set of microbial species. We demonstrate reliable phenotype assignment for Traitar to bacterial genomes from 572 species of 8 phyla, also based on incomplete single-cell genomes and simulated draft genomes. We also showcase its application in metagenomics by verifying and complementing a manual metabolic reconstruction of two novel Clostridiales species based on draft genomes recovered from commercial biogas reactors. Traitar is available at https://github.com/hzi-bifo/traitar.

Bioinformatics

PCR-free library preparation greatly reduces stutter noise at short tandem repeats

Over the past several decades, the forensic and population genetic communities have increasingly leveraged short tandem repeats (STRs) for a variety of applications. The advent of next-generation sequencing technologies and STR-specific bioninformatic tools has enabled the profiling of hundreds of thousands of STRs across the genome. Nonetheless, these genotypes remain error-prone, hindering their utility in downstream analyses. One of the primary drivers of STR genotyping errors are \"stutter\" artifacts arising during the PCR amplification step of library preparation that add or delete copies of the repeat unit in observed sequencing reads. Recently, Illumina developed the TruSeq PCR-free library preparation protocol which eliminates the PCR step and theoretically should reduce stutter error. Here, I compare two high coverage whole genome sequencing datasets prepared with and without the PCR-free protocol. I find that this protocol reduces the percent of reads due to stutter by more than four-fold and results in higher confidence STR genotypes. Notably, stutter at homopolymers was decreased by more than 6fold, making these previously inaccessible loci amenable to STR calling. This technological improvement shows good promise for significantly increasing the feasibility of obtaining high quality STR genotypes from next-generation sequencing technologies.

Bioinformatics

RNA-seq based analysis of population structure within the maize inbred B73

B73 is a variety of maize (Zea mays ssp. mays) widely used in genetic, genomic, and phenotypic research around the world. B73 was also served as the reference genotype for the original maize genome sequencing project. The advent of large-scale RNA-sequencing as a method of measuring gene expression presents a unique opportunity to assess the level of relatedness among individuals identified as variety B73. The level of haplotype conservation and divergence across the genome were assessed using 27 RNA-seq data sets from 20 independent research groups in three countries. Several clearly distinct clades were identified among putatively B73 samples. A number of these blocks were defined by the presence of clearly defined genomic blocks containing a haplotype which did not match the published B73 reference genome. In a number of cases the relationship among B73 samples generated by different research groups recapitulated mentor/mentee relationships within the maize genetics community. A number of regions with distinct, dissimilar, haplotypes were identified in our study. However, when considering the age of the B73 accession - greater than 40 years - and the challenges of maintaining isogenic lines of a naturally outcrossing species, a strikingly high overall level of conservation was exhibited among B73 samples from around the globe.

Bioinformatics

Synlet: an R package for systemically analyzing synthetic lethal RNA interference screen data

SummaryHigh-throughput synthetic lethal RNA interference (RNAi) screen experiments shed important insights on the filed of cancer researches and drug discovery, but a comprehensive software for analyzing the data was not available yet. We present synlet, an R package provided a complete pipeline to process the synthetic lethal RNAi screens data. Synlet provides several methods to access the screen quality, including Z factor and data visualization. B-score and fraction of control or samples normalization methods are implemented in the package. More importantly, synlet facilitates the process of hits selection by implementing several algorithms, providing the possibility to identify high confidence targets.\n\nAvailabilityThe source code is freely available in Bioconductor (http://bioconductor.org/).\n\nContactc.shao@Dkfz-Heidelberg.de

Bioinformatics

GenVisR: Genomic Visualizations in R

SummaryVisualizing and summarizing data from genomic studies continues to be a challenge. Here we introduce the GenVisR package to addresses this challenge by providing highly customizable, publication-quality graphics focused on cohort level genome analyses. GenVisR provides a rapid and easy-to-use suite of genomic visualization tools, while maintaining a high degree of flexibility by leveraging the abilities of ggplot2 and bioconductor.\n\nAvailability and ImplementationGenVisR is an R package available via bioconductor (https://bioconductor.org/packages/GenVisR) under GPLv3. Support is available via GitHub (https://github.com/griffithlab/GenVisR/issues) and the Bioconductor support website.\n\nContactogriffit@genome.wustl.edu, mgriffit@genome.wustl.edu

Bioinformatics

LoLoPicker: Detecting Low Allelic-Fraction Variants in Low-Quality Cancer Samples from Whole-exome Sequencing Data

SummaryWe developed an efficient tool dedicated to call somatic variants from next generation sequencing (NGS) data with the help of a user-defined control panel of non-cancer samples. Compared with other methods, we showed superior performance of LoLoPicker with significantly improved specificity. The algorithm of LoLoPicker is particularly useful for calling low allelic-fraction variants from low-quality cancer samples such as formalin-fixed and paraffin-embedded (FFPE) samples.\n\nImplementation and AvailabilityThe main scripts are implemented in Python 2.7.8 and the package is released at https://github.com/jcarrotzhang/LoLoPicker.

Bioinformatics