Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Assessment of Circulating Copy Number Variant Detection for Cancer Screening

Current high-sensitivity cancer screening methods suffer from false positive rates that lead to numerous unnecessary procedures and questionable public health benefit overall. Detection of circulating tumor DNA (ctDNA) has the potential to transform cancer screening. Thus far, nearly all ctDNA studies have focused on detection of tumor-specific point mutations. However, ctDNA point mutation detection methods developed to date lack either the scope or sensitivity necessary to be useful for cancer screening, due to the extremely low (<1%) ctDNA fraction derived from early stage tumors. We suggest that tumor-derived copy number variant (CNV) detection is theoretically a superior means of ctDNA-based cancer screening for many tumor types, given that, relative to point mutations, each individual tumor CNV contributes a much larger number of ctDNA fragments to the overall pool of circulating DNA. Here we perform an in silico assessment of the potential for ctDNA CNV-based cancer screening across many common cancers.

Bioinformatics

Vcfanno: fast, flexible annotation of genetic variants

BackgroundThe integration of genome annotations and reference databases is critical to the identification of genetic variants that may be of interest in studies of disease or other traits. However, comprehensive variant annotation with diverse file formats is difficult with existing methods.\n\nResultsWe have developed vcfanno as a flexible toolset that simplifies the annotation of genetic variants in VCF format. Vcfanno can extract and summarize multiple attributes from one or more annotation files and append the resulting annotations to the INFO field of the original VCF file. Vcfanno also integrates the lua scripting language so that users can easily develop custom annotations and metrics. By leveraging a new parallel \"chromosome sweeping\" algorithm, it enables rapid annotation of both whole-exome and whole-genome datasets. We demonstrate this performance by annotating over 85.3 million variants in less than 17 minutes (>85,000 variants per second) with 50 attributes from 17 commonly used genome annotation resources.\n\nConclusionsVcfanno is a flexible software package that provides researchers with the ability to annotate genetic variation with a wide range of datasets and reference databases in diverse genomic formats.\n\nAvailabilityThe vcfanno source code is available at https://github.com/brentp/vcfanno under the MIT license, and platform-specific binaries are available at https://github.com/brentp/vcfanno/releases. Detailed documentation is available at http://brentp.github.io/vcfanno/, and the code underlying the analyses presented can be found at https://github.com/brentp/vcfanno/tree/master/scripts/paper.

Bioinformatics

OncoScape: Exploring the cancer aberration landscape by genomic data fusion

Although large-scale efforts for molecular profiling of cancer samples provide multiple data types for many samples, most approaches for finding candidate cancer genes rely on somatic mutations and DNA copy number only. We present a new method, OncoScape, which, for the first time, exploits five complementary data types across 11 cancer types to identify new candidate cancer genes. We find many rarely mutated genes that are strongly affected by other aberrations. We retrieve the majority of known cancer genes but also new candidates such as STK31 and MSRA with very high confidence. Several genes show a dual oncogene-and tumor suppressor-like behavior depending on the tumor type. Most notably, the well-known tumor suppressor RB1 shows strong oncogene-like signal in colon cancer. We applied OncoScape to cell lines representing ten cancer types, providing the most comprehensive comparison of aberrations in cell lines and tumor samples to date. This revealed that glioblastoma, breast and colon cancer show strong similarity between cell lines and tumors, while head and neck squamous cell carcinoma and bladder cancer, exhibit very little similarity between cell lines and tumors. To facilitate exploration of the cancer aberration landscape, we created a web portal enabling interactive analysis of OncoScape results.

Bioinformatics

A simple analytical formula to compute the residual Mutual Information between pairs of data vectors

SummaryThe Mutual Information of pairs of data vectors, for example sequence alignment positions or gene expression profiles, is a quantitative measure of the interdependence between the data. However, data vectors based on a finite number of samples retain non-zero Mutual Information values even for completely random data, which is referred to as background or residual Mutual Information. Estimates of the residual Mutual Information have so far been obtained through heuristic or numerical approximations. Here we introduce a simple analytical formula for the computation of the residual Mutual Information that yields precise values and does not require the joint probabilities between the vector elements as input.\n\nAvailability and ImplementationA C program arMI is available at http://mathbio.crick.ac.uk/wiki/Software#arMI. Using an input alignment in FASTA format or alternatively an internally created random alignment of specified length and depth, the program computes three types of Mutual information: (i) Shannons Mutual Information between all pairs of alignment columns; (ii) the numerical residual Mutual Information by using the same formula on the randomised (shuffled) data; (iii) the analytical residual Mutual Information introduced here. The package depends on the GNU Scientific Library, which is used for vector and matrix operations, factorial expressions and random number generation (Galassi et al., 2009). Reference alignments and result data are included in the program package in the folder tests. The R environment was used for statistics and plotting (R Core Team, 2014).\n\nContactJens.Kleinjung@crick.ac.uk\n\nSupplementary MaterialA detailed derivation of the analytical formula is given in the Supplementary Material.

Bioinformatics

Phylogeny-aware Identification and Correction of Taxonomically Mislabeled Sequences

Molecular sequences in public databases are mostly annotated by the submitting authors without further validation. This procedure can generate erroneous taxonomic sequence labels. Mislabeled sequences are hard to identify, and they can induce downstream errors because new sequences are typically annotated using existing ones. Furthermore, taxonomic mislabelings in reference sequence databases can bias metagenetic studies which rely on the taxonomy. Despite significant efforts to improve the quality of taxonomic annotations, the curation rate is low because of the labour-intensive manual curation process.\n\nHere, we present SATIVA, a phylogeny-aware method to automatically identify taxonomically mislabeled sequences ( mislabels) using statistical models of evolution. We use the Evolutionary Placement Algorithm (EPA) to detect and score sequences whose taxonomic annotation is not supported by the underlying phylogenetic signal, and automatically propose a corrected taxonomic classification for those. Using simulated data, we show that our method attains high accuracy for identification (96.9% sensitivity / 91.7% precision) as well as correction (94.9% sensitivity / 89.9% precision) of mislabels. Furthermore, an analysis of four widely used microbial 16S reference databases (Greengenes, LTP, RDP and SILVA) indicates that they currently contain between 0.2% and 2.5% mislabels. Finally, we use SATIVA to perform an in-depth evaluation of alternative taxonomies for Cyanobacteria.\n\nSATIVA is freely available at https://github.com/amkozlov/sativa.

Bioinformatics

The Ensembl Variant Effect Predictor

The Ensembl Variant Effect Predictor (VEP) is a powerful toolset for the analysis, annotation and prioritization of genomic variants, including in non-coding regions.\n\nThe VEP accurately predicts the effects of sequence variants on transcripts, protein products, regulatory regions and binding motifs by leveraging the high quality, broad scope, and integrated nature of the Ensembl databases. In addition, it enables comparison with a large collection of existing publicly available variation data within Ensembl to provide insights into population and ancestral genetics, phenotypes and disease.\n\nThe VEP is open source and free to use. It is available via a simple web interface (http://www.ensembl.org/vep), a powerful downloadable package, and both Ensembls Perl and REST application program interface (API) services.

Bioinformatics

Exploring miRNAs as the key to understand symptoms induced by ZIKA virus infection through a collaborative database.

Zika virus (ZIKV) is an emerging mosquito-borne flavivirus, first isolated in 1947 from the serum of a pyrexial rhesus monkey caged in the Zika Forest (Uganda/Africa)1. In 2007 ZIKV was reported to of been responsible for an outbreak of relatively mild disease, characterized by rash, arthralgia, and conjunctivitis on Yap Island, in the western Pacific Ocean2. In the past year, ZIKV has been circulating in the Americas, probably introduced through Easter Island (Chile), by French Polynesians3. In early 2015, a new outbreak was recognized in northeast Brazil4, where concerns over its possible links with infant microcephaly have been discussed5,6. Providing a definitive link between ZIKV infection and birth defects is still a big challenge7. MicroRNAs (miRNAs), are small noncoding RNAs that regulating post-transcriptional gene expression by translational repression, and play important roles in viral pathogenesis8 and brain development9. It is estimated that more than 60% of human protein-coding genes contain at least one conserved miRNA-binding site10. The potential for flavivirus-mediated miRNA signaling dysfunction in brain-tissue develop provides a compelling mechanism underlying perceived linked between ZIKV and microcephaly. Here, we provide strong evidences toward to understand the mechanism in which miRNAs can be linked to the \"congenital Zika syndrome\" symptoms. Moreover, following World Health Organization (WHO) recommendations11, we have assembled a database that could help target mechanistic investigations of this possible relationship between ZIKV symptoms and miRNA mediated human gene expression, helping to foster potential targets for therapy.

Bioinformatics

Combining multiple tools outperforms individual methods in gene set enrichment analyses

Gene set enrichment (GSE) analysis allows researchers to efficiently extract biological insight from long lists of differentially expressed genes by interrogating them at a systems level. In recent years, there has been a proliferation of GSE analysis methods and hence it has become increasingly difficult for researchers to select an optimal GSE tool based on their particular data set. Moreover, the majority of GSE analysis methods do not allow researchers to simultaneously compare gene set level results between multiple experimental conditions.\n\nResults: The ensemble of genes set enrichment analyses (EGSEA) is a method developed for RNA-sequencing data that combines results from twelve algorithms and calculates collective gene set scores to improve the biological relevance of the highest ranked gene sets. redEGSEAs gene set database contains around 25,000 gene sets from sixteen collections. It has multiple visualization capabilities that allow researchers to view gene sets at various levels of granularity. EGSEA has been tested on simulated data and on a number of human and mouse data sets and, based on biologists' feedback, consistently outperforms the individual tools that have been combined. Our evaluation demonstrates the superiority of the ensemble approach for GSE analysis, and its utility to effectively and efficiently extrapolate biological functions and potential involvement in disease processes from lists of differentially regulated genes.\n\nAvailability and Implementation: EGSEA is available as an R package at http://www.bioconductor.org/packages/EGSEA/. The gene sets collections are available in the R package EGSEAdata from http://www.bioconductor.org/packages/EGSEA/.

Bioinformatics

mLDM: a new hierarchical Bayesian statistical model for sparse microbioal association discovery

Interpretive analysis of metagenomic data depends on an understanding of the underlying associations among microbes from metagenomic samples. Although several statistical tools have been developed for metage-nomic association studies, they suffer from compositional bias or fail to take into account environmental factors that directly affect the composition of a given microbial community. In this paper, we propose metagenomic Lognormal-Dirichlet-Multinomial (mLDM), a hierarchical Bayesian model with sparsity constraints to bypass compositional bias and discover new associations among microbes and between microbes and environmental factors. The mLD-M model can 1) infer both conditionally dependent associations among microbes and direct associations between microbes and environmental factors; 2) consider both compositional bias and variance of metagenomic data; and 3) estimate absolute abundance for microbes. Thus, conditionally dependent association can capture direct relationship underlying microbial pairs and remove the indirect connections induced from other common factors. Empirical studies show the effectiveness of the mLDM model, using both synthetic data and the TARA Oceans eukaryotic data by comparing it with several state-of-the-art methodologies. Finally, mLDM is applied to western English Channel data and finds some interesting associations.

Bioinformatics

AnnoLnc: a web server for systematically annotating novel human lncRNAs

Although the repertoire of human lncRNAs has rapidly expanded, their biological function and regulation remain largely elusive. Here, we present AnnoLnc (http://annolnc.cbi.pku.edu.cn), an online portal for systematically annotating newly identified human lncRNAs. AnnoLnc offers a full spectrum of annotations covering genomic location, RNA secondary structure, expression, transcriptional regulation, miRNA interaction, protein interaction, genetic association and evolution, as well as an abstraction-based text summary and various intuitive figures to help biologists quickly grasp the essentials. In addition to an intuitive and mobile-friendly Web interactive design, AnnoLnc supports batch analysis and provides JSON-based Web Service APIs for programmatic analysis. To the best of our knowledge, AnnoLnc is the first web server to provide on-the-fly and systematic annotation for newly identified human lncRNAs. Some case studies have shown the power of AnnoLnc to inspire novel hypotheses.

Bioinformatics

Discovery of Primary, Cofactor, and Novel Transcription Factor Binding Site Motifs by Recursive, Thresholded Entropy Minimization

Data from ChIP-seq experiments can determine the genome-wide binding specificities of transcription factors (TFs) and other regulatory proteins. In the present study, we analyzed 745 ENCODE ChIP-seq peak datasets of 189 human TFs with a novel motif discovery method that is based on recursive, thresholded entropy minimization. This method is able to distinguish correct information models from noisy motifs, quantify the strengths of individual sites based on affinity, and detect adjacent cofactor binding sites that coordinate with primary TFs. We derived homogeneous and bipartite information models for 89 sequence-specific TFs, which enabled discovery of 24 cofactor motifs for 118 TFs, and revealed 6 high-confidence novel motifs. The reliability and accuracy of these models were determined via three independent quality control criteria, including the detection of experimentally proven binding sites, comparison with previously published motifs and statistical analyses. We also predict previously unreported TF cobinding interactions, and new components of known TF complexes. Because they are based on information theory, the derived models constitute a powerful tool for detecting and predicting the effects of variants in known binding sites, and predicting previously unrecognized binding sites and target genes.

Bioinformatics

A fast and accurate method for detection of IBD shared haplotypes in genome-wide SNP data

Identical by descent (IBD) segments are used to understand a number of fundamental issues in genetics. IBD segments are typically detected using long stretches of identical alleles between haplotypes in whole-genome SNP data. Phase or SNP call errors in genomic data can degrade accuracy of IBD detection and lead to false positive calls, false negative calls, and under- or overextension of true IBD segments. Furthermore, the number of comparisons increases quadratically with sample size, requiring high computational efficiency. We developed a new IBD segment detection program, FISHR (Find IBD Shared Haplotypes Rapidly), in an attempt to accurately detect IBD segments and to better estimate their endpoints using an algorithm that is fast enough to be deployed on the very large whole-genome SNP datasets. We compared the performance of FISHR to three leading IBD segment detection programs: GERMLINE, refinedIBD, and HaploScore. Using simulated and real genomic sequence data, we show that FISHR is slightly more accurate than all programs at detecting long (>3 cM) IBD segments but slightly less accurate than refinedIBD at detecting short (~1 cM) IBD segments. Moreover, FISHR outperforms all programs in determining the true endpoints of IBD segments, which is important for several reasons. FISHR takes two to four times longer than GERMLINE to run, whereas both GERMLINE and FISHR were orders of magnitude faster than refinedIBD and HaploScore. Overall, FISHR provides accurate IBD detection in unrelated individuals and is computationally efficient enough to be utilized on large SNP datasets > 20,000 individuals.

Bioinformatics

Crunch: Completely Automated Analysis of ChIP-seq Data

Although it has become routine for experimental groups to apply ChIP-seq technology to quantitatively characterize the genome-wide binding of transcription factors (TFs), computational analysis procedures remain far from standardized, making it difficult to meaningfully compare ChIP-seq results across experiments. In addition, while genome-wide binding patterns must ultimately be determined by local constellations of binding sites in the DNA, current analysis is typically limited to a standard search for enriched motifs in ChIP-seq peaks.\n\nHere we present Crunch, a completely automated computational method that performs all ChIP-seq analysis from quality control through read mapping and peak detecting, and integrates comprehensive modeling of the ChIP signal in terms of known and novel binding motifs, quantifying the contribution of each motif, and annotating which combinations of motifs explain each binding peak.\n\nApplying Crunch to 128 ChIP-seq datasets from the ENCODE project we find that TFs naturally separate into solitary TFs, for which a single motif explains the ChIP-peaks, and co-binding TFs for which multiple motifs co-occur within peaks. Moreover, for most datasets the motifs that Crunch identified de novo outperform known motifs and both the set of co-binding motifs and the top motif of solitary TFs are consistent across experiments and cell lines. Crunch is implemented as a web server (crunch.unibas.ch), enabling standardized analysis of any collection of ChIP-seq datasets by simply uploading raw sequencing data. Results are provided both in a graphical interface and as downloadable files.

Bioinformatics

Highly accessible AU-rich regions in 3′ untranslated regions are hotspots for binding of proteins and miRNAs

BackgroundMicroRNAs (miRNAs) are endogenous short non-coding RNAs involved in the regulation of gene expression at the post-transcriptional level typically by promoting destabilization or translational repression of target RNAs. Sometimes this regulation is absent or different, which likely is the result of interactions with other post-transcriptional factors, particularly RNA-binding proteins (RBPs). Despite the importance of the interactions between RBPs and miRNAs, little is known about how they affect post-transcriptional regulation in a global scale.\n\nResultsIn this study, we have analyzed CLIP datasets of 49 RBPs in HEK293 cells with the aim of understanding the interplay between RBPs and miRNAs in post-transcriptional regulation. Our results show that RBPs bind preferentially in conserved regulatory hotspots that frequently contain miRNA target sites. This organization facilitates the competition and cooperation among RBPs and the regulation of miRNA target site accessibility. In some cases RBP enrichment on target sites correlates with miRNA expression, suggesting coordination between the regulatory factors. However, in most cases, competition among factors is the most plausible interpretation of our data. Upon AGO2 knockdown, transcripts that contain such hotspots that overlap target sites of expressed miRNAs in 3UTRs are significantly less up-regulated than transcripts without them, suggesting that RBP binding limits miRNA accessibility.\n\nConclusionsWe show that RBP binding is concentrated in regulatory hotspots in 3UTRs. The presence of these hotspots facilitates the interaction among post-transcriptional regulators, that interact or compete with each other under different conditions. These hotspots are enriched in genes with regulatory functions such as DNA binding and RNA binding. Taken together, our results suggest that hotspots are important regulatory regions that define an extra layer of auto-regulatory control of post-transcriptional regulation.

Bioinformatics

A method for downstream analysis of gene set enrichment results facilitates the biological interpretation of vaccine efficacy studies

Gene set enrichment analysis (GSEA) is a widely employed method for analyzing gene expression profiles. The approach uses annotated sets of genes, identifies those that are coordinately up- or down-regulated in a biological comparison of interest, and thereby elucidates underlying biological processes relevant to the comparison. As the number of gene sets available in various collections for enrichment analysis has grown, the resulting lists of significant differentially regulated gene sets may also become larger, leading to the need for additional downstream analysis of GSEA results. Here we present a method that allows the rapid identification of a small number of co-regulated groups of genes - \"leading edge metagenes\" (LEMs) - from high scoring sets in GSEA results. LEM are sub-signatures which are common to multiple gene sets and that \"explain\" their enrichment specific to the experimental dataset of interest. We show that LEMs contain more refined lists of context-dependent and biologically meaningful genes than the parental gene sets. LEM analysis of the human vaccine response using a large database of immune signatures identified core biological processes induced by five different vaccines in datasets from human peripheral blood mononuclear cells (PBMC). Further study of these biological processes over time following vaccination showed that at day 3 post-vaccination, vaccines derived from viruses or viral subunits exhibit patterns of biological processes that are distinct from protein conjugate vaccines; however, by day 7 these differences were less pronounced. This suggests that the immune response to diverse vaccines eventually converge to a common transcriptional response. LEM analysis can significantly reduce the dimensionality of enriched gene sets, improve the identification of core biological processes active in a comparison of interest, and simplify the biological interpretation of GSEA results.\n\nAuthor SummaryGenome-wide expression profiling is a widely used tool to identify biological mechanisms in a comparison of interest. One analytic method, Gene set enrichment analysis (GSEA) uses annotated sets of genes and identifies those that are coordinately up- or down-regulated in a biological comparison of interest. This approach capitalizes on the fact that alternations in biological processes often cause the coordinated change of a large number of genes. However, as the number of gene sets available in various collections for enrichment analysis has grown, the resulting lists of significant differentially regulated gene sets may also become larger, leading to the need for additional downstream analysis of GSEA results. Here we present a method that allows the identification of a small number of co-regulated groups of genes - \"leading edge metagenes\" (LEMs) - from high scoring sets in GSEA results. We show that LEMs contain more refined lists of context-dependent biologically meaningful genes than the parental gene sets and demonstrate the utility of this approach in analyzing the transcriptional response to vaccination. LEM analysis can significantly reduce the dimensionality of enriched gene sets, improve the identification of core biological processes active in a comparison of interest, and facilitate the biological interpretation of GSEA results.

Bioinformatics

RSQ: a statistical method for quantification of isoform-specific structurome using transcriptome-wide structural profiling data

The structure of RNA, which is considered to be a second layer of information alongside the genetic code, provides fundamental insights into the cellular function of both coding and non-coding RNAs. Several high-throughput technologies have been developed to profile transcriptome-wide RNA structures, i.e., the structurome. However, it is challenging to interpret the profiling data because the observed data represent an average over different RNA conformations and isoforms with different abundance. To address this challenge, we developed an RNA structurome quantification method (RSQ) to statistically model the distribution of reads over both isoforms and RNA conformations, and thus provide accurate quantification of the isoform-specific structurome. The quantified RNA structurome enables the comparison of isoform-specific conformations between different conditions, the exploration of RNA conformation variation affected by single nucleotide polymorphism (SNP),and the measurement of RNA accessibility for binding of either small RNAs in RNAi-based assays or RNA binding protein in transcriptional regulation. The model used in our method sheds new light on the potential impact of the RNA structurome on gene regulation.

Bioinformatics

From genomes to phenotypes: Traitar, the microbial trait analyzer

The number of sequenced genomes is growing exponentially, profoundly shifting the bottleneck from data generation to genome interpretation. Traits are often used to characterize and distinguish bacteria, and are likely a driving factor in microbial community composition, yet little is known about the traits of most microbes. We describe Traitar, the microbial trait analyzer, which is a fully automated software package for deriving phenotypes from the genome sequence. Traitar provides phenotype classifiers to predict 67 traits related to the use of various substrates as carbon and energy sources, oxygen requirement, morphology, antibiotic susceptibility, proteolysis and enzymatic activities. Furthermore, it suggests protein families associated with the presence of particular phenotypes. Our method uses L1-regularized L2-loss support vector machines for phenotype assignments based on phyletic patterns of protein families and their evolutionary histories across a diverse set of microbial species. We demonstrate reliable phenotype assignment for Traitar to bacterial genomes from 572 species of 8 phyla, also based on incomplete single-cell genomes and simulated draft genomes. We also showcase its application in metagenomics by verifying and complementing a manual metabolic reconstruction of two novel Clostridiales species based on draft genomes recovered from commercial biogas reactors. Traitar is available at https://github.com/hzi-bifo/traitar.

Bioinformatics

PCR-free library preparation greatly reduces stutter noise at short tandem repeats

Over the past several decades, the forensic and population genetic communities have increasingly leveraged short tandem repeats (STRs) for a variety of applications. The advent of next-generation sequencing technologies and STR-specific bioninformatic tools has enabled the profiling of hundreds of thousands of STRs across the genome. Nonetheless, these genotypes remain error-prone, hindering their utility in downstream analyses. One of the primary drivers of STR genotyping errors are \"stutter\" artifacts arising during the PCR amplification step of library preparation that add or delete copies of the repeat unit in observed sequencing reads. Recently, Illumina developed the TruSeq PCR-free library preparation protocol which eliminates the PCR step and theoretically should reduce stutter error. Here, I compare two high coverage whole genome sequencing datasets prepared with and without the PCR-free protocol. I find that this protocol reduces the percent of reads due to stutter by more than four-fold and results in higher confidence STR genotypes. Notably, stutter at homopolymers was decreased by more than 6fold, making these previously inaccessible loci amenable to STR calling. This technological improvement shows good promise for significantly increasing the feasibility of obtaining high quality STR genotypes from next-generation sequencing technologies.

Bioinformatics