Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,423 records · Page 79Linked to original sources

RNA-seq based analysis of population structure within the maize inbred B73

B73 is a variety of maize (Zea mays ssp. mays) widely used in genetic, genomic, and phenotypic research around the world. B73 was also served as the reference genotype for the original maize genome sequencing project. The advent of large-scale RNA-sequencing as a method of measuring gene expression presents a unique opportunity to assess the level of relatedness among individuals identified as variety B73. The level of haplotype conservation and divergence across the genome were assessed using 27 RNA-seq data sets from 20 independent research groups in three countries. Several clearly distinct clades were identified among putatively B73 samples. A number of these blocks were defined by the presence of clearly defined genomic blocks containing a haplotype which did not match the published B73 reference genome. In a number of cases the relationship among B73 samples generated by different research groups recapitulated mentor/mentee relationships within the maize genetics community. A number of regions with distinct, dissimilar, haplotypes were identified in our study. However, when considering the age of the B73 accession - greater than 40 years - and the challenges of maintaining isogenic lines of a naturally outcrossing species, a strikingly high overall level of conservation was exhibited among B73 samples from around the globe.

Bioinformatics

Synlet: an R package for systemically analyzing synthetic lethal RNA interference screen data

SummaryHigh-throughput synthetic lethal RNA interference (RNAi) screen experiments shed important insights on the filed of cancer researches and drug discovery, but a comprehensive software for analyzing the data was not available yet. We present synlet, an R package provided a complete pipeline to process the synthetic lethal RNAi screens data. Synlet provides several methods to access the screen quality, including Z factor and data visualization. B-score and fraction of control or samples normalization methods are implemented in the package. More importantly, synlet facilitates the process of hits selection by implementing several algorithms, providing the possibility to identify high confidence targets.\n\nAvailabilityThe source code is freely available in Bioconductor (http://bioconductor.org/).\n\nContactc.shao@Dkfz-Heidelberg.de

Bioinformatics

GenVisR: Genomic Visualizations in R

SummaryVisualizing and summarizing data from genomic studies continues to be a challenge. Here we introduce the GenVisR package to addresses this challenge by providing highly customizable, publication-quality graphics focused on cohort level genome analyses. GenVisR provides a rapid and easy-to-use suite of genomic visualization tools, while maintaining a high degree of flexibility by leveraging the abilities of ggplot2 and bioconductor.\n\nAvailability and ImplementationGenVisR is an R package available via bioconductor (https://bioconductor.org/packages/GenVisR) under GPLv3. Support is available via GitHub (https://github.com/griffithlab/GenVisR/issues) and the Bioconductor support website.\n\nContactogriffit@genome.wustl.edu, mgriffit@genome.wustl.edu

Bioinformatics

LoLoPicker: Detecting Low Allelic-Fraction Variants in Low-Quality Cancer Samples from Whole-exome Sequencing Data

SummaryWe developed an efficient tool dedicated to call somatic variants from next generation sequencing (NGS) data with the help of a user-defined control panel of non-cancer samples. Compared with other methods, we showed superior performance of LoLoPicker with significantly improved specificity. The algorithm of LoLoPicker is particularly useful for calling low allelic-fraction variants from low-quality cancer samples such as formalin-fixed and paraffin-embedded (FFPE) samples.\n\nImplementation and AvailabilityThe main scripts are implemented in Python 2.7.8 and the package is released at https://github.com/jcarrotzhang/LoLoPicker.

Bioinformatics

Pathway-Structured Predictive Model for Cancer Survival Prediction: A Two-Stage Approach

Heterogeneity in terms of tumor characteristics, prognosis, and survival among cancer patients has been a persistent problem for many decades. Currently, prognosis and outcome predictions are made based on clinical factors and/or by incorporating molecular profiling data. However, inaccurate prognosis and prediction may result by using only clinical or molecular information directly. One of the main shortcomings of past studies is the failure to incorporate prior biological information into the predictive model, given strong evidence of pathway-based genetic nature of cancer, i.e. the potential for oncogenes to be grouped into pathways based on biological functions such as cell survival, proliferation and metastatic dissemination.\n\nTo address this problem, we propose a two-stage procedure to incorporate pathway information into the prognostic modeling using large-scale gene expression data. In the first stage, we fit all predictors within each pathway using penalized Cox model (Lasso, Ridge and Elastic Net) and Bayesian hierarchical Cox model. In the second stage, we combine the cross-validated prognostic scores of all pathways obtained in the first stage as new predictors to build an integrated prognostic model for prediction. We apply the proposed method to analyze breast cancer data from The Cancer Genome Atlas (TCGA), predicting overall survival using clinical data and gene expression profiling. The data includes ~20000 genes mapped into 109 pathways for 505 patients. The results show that the proposed approach not only improves survival prediction compared with the alternative analysis that ignores the pathway information, but also identifies significant biological pathways.

Bioinformatics

Prediction of kinase-specific phosphorylation sites through an integrative model of protein context and sequence

The identification of kinase substrates and the specific phosphorylation sites they regulate is an important factor in understanding protein function regulation and signalling pathways. Computational prediction of kinase targets - assigning kinases to putative substrates, and selecting from protein sequence the sites that kinases can phosphorylate - requires the consideration of both the cellular context that kinases operate in, as well as their binding affinity. This consideration enables investigation of how phosphorylation influences a range of biological processes.\n\nWe report here a novel probabilistic model for the classification of kinase-specific phos-phorylation sites from sequence across three model organisms: human, mouse and yeast. The model incorporates position-specific amino acid frequencies, and counts of co-occurring amino acids from kinase binding sites in a kinase- and family-specific manner. We show how this model can be seamlessly integrated with protein interactions and cell-cycle abundance profiles. When evaluating the prediction accuracy of our method, PhosphoPICK, on an independent hold-out set of kinase-specific phosphorylation sites, we found it achieved an average specificity of 97% while correctly predicting 32% of true positives. We also compared PhosphoPICKs ability, through cross-validation, to predict kinase-specific phosphorylation sites with alternative methods, and found that at high levels of specificity PhosphoPICK outperforms alternative methods for most comparisons made.\n\nWe investigated the relationship between experimentally confirmed phosphorylation sites and predicted nuclear localisation signals by predicting the most likely kinases to be regulating the phosphorylated residues immediately upstream or downstream from the localisation signal. We show that kinases PKA, Akt1 and AurB have an over-representation of predicted binding sites at particular positions downstream from predicted nuclear localisation signals, indicating a greater role for phosphorylation in regulating the nuclear import of proteins than previously thought.\n\nPhosphoPICK is freely available online as a web-service at http://bioinf.scmb.uq.edu.au/phosphopick.\n\nAbbreviations

Bioinformatics

Modeling methyl-sensitive transcription factor motifs with an expanded epigenetic alphabet

Transcription factors bind DNA in specific sequence contexts. In addition to distinguishing one nucleobase from another, some transcription factors can distinguish between unmodified and modified bases. Current models of transcription factor binding tend not take DNA modifications into account, while the recent few that do often have limitations. This makes a comprehensive and accurate profiling of transcription factor affinities difficult. Here, we developed methods to identify transcription factor binding sites in modified DNA. Our models expand the standard A/C/G/T DNA alphabet to include cytosine modifications. We developed Cytomod to create modified genomic sequences and enhanced the Multiple EM for Motif Elicitation (MEME) Suite by adding the capacity to handle custom alphabets. We adapted the well-established position weight matrix (PWM) model of transcription factor binding affinity to this expanded DNA alphabet. Using these methods, we identified modification-sensitive transcription factor binding motifs. We confirmed established binding preferences, such as the preference of ZFP57 and C/EBP{beta} for methylated motifs and the preference of c-Myc for unmethylated E-box motifs. Using known binding preferences to tune model parameters, we discovered novel modified motifs for a wide array of transcription factors. Finally, we validated predicted binding preferences of OCT4 using cleavage under targets and release using nuclease (CUT&RUN) experiments across conventional, methylation-, and hydroxymethylation-enriched sequences. Our approach readily extends to other DNA modifications. As more genome-wide single-base resolution modification data becomes available, we expect that our method will yield insights into altered transcription factor binding affinities across many different modifications.

Bioinformatics

Strawberry: fast and accurate genome-guided transcript reconstruction and quantification from RNA-seq

We propose a novel method and computational tool, Strawberry, for transcript reconstruction and quantification from paired-end RNA-seq data under the guidance of genome alignment and independent of gene annotation. Strawberry achieves this through disentangling assembly and quantification in a sequential manner. The application of a fast flow network algorithm for assembly speeds up the construction of a parsimonious set of transcripts. The resulting reduced data representation improves the efficiency of expression-level quantification. Strawberry leverages the speed and accuracy of transcript assembly and quantification in such a way that processing 10 million simulated reads (after alignment) requires only 90 seconds using a single thread while achieving over 92% correlation with the ground truth, making it the state-of-the-art method. Strawberry outperforms Cufflinks and StringTie, the two other leading methods, in many aspects, including the number of corrected assembled transcripts and the correlation with the ground truth of simulated RNA-seq data. Availability: Strawberry is written in C++11, and is available as open source software at https://github.com/ruolin/Strawberry under the GPLv3 license.

Bioinformatics

Finding De novo methylated DNA motifs

Increasing evidence has shown that posttranslational modifications (PTMs) such as methylation and hydroxymethylation on cytosine would greatly impact the binding of transcription factors (TFs). However, there is a lack of motif finding algorithms with the function to search for motifs with PTMs. In this study, we expend on our previous motif finding pipeline Epigram to provide systematic de novo motif discovery and performance evaluation on methylated DNA motifs. Using the tool, we were able to identified methylated motifs in Arabidopsis DAP-seq data that were previously demonstrated to contain such motifs1. When applied to TF ChIP-seq and DNA methylome data in H1 and GM12878, our method successfully identified novel methylated motifs that can be recognized by the TFs or their co-factors. We also observed spacing constraint between the canonical motif of the TF and the newly discovered methylated motifs, which suggests operative recognition of these cis-elements by collaborative proteins.

Bioinformatics

Hierarchical Association Coefficient Algorithm

Suppose that members in a universal set categorized based on observations, and that categories can be stratified based on the average of observations within each category. Two sorting extremes can be obtained from the perspective of arbitrariness of an order of observations. The first sorting extreme is an increasing order of observations on ascendingly stratified categories. The second sorting extreme is a decreasing order of observations on ascendingly stratified categories. Hierarchical association coefficient (HA-coefficient) algorithm is based on a principle that any order of observations in stratified categorization can be placed between the two sorting extremes. The algorithm produces a proportion of how much an order of observations in stratified categorization is close to the first sorting extreme, or how much an order of categorized observations is distant from the second sorting extreme. This paper introduces a theory about the HA-coefficient algorithm, and shows its applications with example data. In addition, proving a reliability of the algorithm is shown through a simulation.

Bioinformatics

Improved Placement of Multi-Mapping Small RNAs

High-throughput sequencing of small RNAs (sRNA-seq) is a popular method used to discover and annotate microRNAs (miRNAs), endogenous short interfering RNAs (siRNAs) and Piwi-associated RNAs (piRNAs). One of the key steps in sRNA-seq data analysis is alignment to a reference genome. sRNA-seq libraries often have a high proportion of reads which align to multiple genomic locations, which makes determining their true loci of origin difficult. Commonly used sRNA-seq alignment methods result in either very low precision (choosing an alignment at random) or sensitivity (ignoring multi-mapping reads). Here, we describe and test an sRNA-seq alignment strategy that uses local genomic context to guide decisions on proper placements of multi-mapped sRNA-seq reads. Tests using simulated sRNA-seq data demonstrated that this local-weighting method outperforms other alignment strategies using three different plant genomes. Experimental analyses with real sRNA-seq data also indicate superior performance of local-weighting methods for both plant miRNAs and heterochromatic siRNAs. The local-weighting methods we have developed are implemented as part of the sRNA-seq analysis program ShortStack, which is freely available under a general public license. Improved genome alignments of sRNA-seq data should increase the quality of downstream analyses and genome annotation efforts.\n\nArticle SummaryHigh-throughput sequencing of small RNAs (sRNA-seq) is a frequently used technique in the study of small RNAs. Alignment to a reference genome is a key step in processing sRNA-seq libraries, but suffers from enormous rates of multi-mapping reads. Current methods for sRNA-seq alignment either place these reads randomly or ignore them, both of which distort downstream analyses. Here, we describe a locality-based weighting approach to make better decisions of placement of multi-mapped sRNA-seq data, and test our implementation of this method. We find that our method gives superior performance in terms of placing multi-mapped sRNA-seq data. An implementation of our method is freely available within the ShortStack small RNA analysis program. Use of this method may dramatically improve genome-wide analyses of small RNAs.

Bioinformatics

AVOCADO: Visualization of Workflow-Derived Data Provenance for Reproducible Biomedical Research

A major challenge of data-driven biomedical research lies in the collection and representation of data provenance information to ensure reproducibility of findings. In order to communicate and reproduce multi-step analysis workflows executed on datasets that contain data for dozens or hundreds of samples, it is crucial to be able to visualize the provenance graph at different levels of aggregation. Most existing approaches are based on node-link diagrams, which do not scale to the complexity of typical data provenance graphs. In our proposed approach we reduce the complexity of the graph using hierarchical and motif-based aggregation. Based on user action and graph attributes a modular degree-of-interest (DoI) function is applied to expand parts of the graph that are relevant to the user. This interest-driven adaptive provenance visualization approach allows users to review and communicate complex multi-step analyses, which can be based on hundreds of files that are processed by numerous workflows. We integrate our approach into an analysis platform that captures extensive data provenance information and demonstrate its effectiveness by means of a biomedical usage scenario.

Bioinformatics

mixMC: a multivariate statistical framework to gain insight into Microbial Communities

Culture independent techniques, such as shotgun metagenomics and 16S rRNA amplicon sequencing have dramatically changed the way we can examine microbial communities. Recently, changes in microbial community structure and dynamics have been associated with a growing list of human diseases. The identification and comparison of bacteria driving those changes requires the development of sound statistical tools, especially if microbial biomarkers are to be used in a clinical setting.\n\nWe present mixMC, a novel multivariate data analysis framework for metagenomic biomarker discovery. mixMC accounts for the compositional nature of 16S data and enables detection of subtle differences when high inter-subject variability is present due to microbial sampling performed repeatedly on the same subjects but in multiple habitats. Through data dimension reduction the multivariate methods provide insightful graphical visualisations to characterise each type of environment in a detailed manner.\n\nWe applied mixMC to 16S microbiome studies focusing on multiple body sites in healthy individuals, compared our results with existing statistical tools and illustrated added value of using multivariate methodologies to fully characterise and compare microbial communities.

Bioinformatics

iFORM: incorporating Find Occurrence of Regulatory Motifs

MotivationAccurately identifying binding sites of transcription factors (TFs) is crucial to understand the mechanisms of transcriptional regulation and human disease.\n\nResultsWe present incorporating Find Occurrence of Regulatory Motifs (iFORM), an easy-to-use tool for scanning DNA sequence with TF motifs described as position weight matrices (PWMs). iFORM achieves higher accuracy and sensitivity by integrating the results from five classical motif discovery programs based on Fishers combined probability test. We have used iFORM to provide accurate results on a variety of data in the ENCODE Project and the NIH Roadmap Epigenomics Project, and has demonstrated its utility to further understand individual roles of functional elements.\n\nAvailabilityiFORM can be freely accessed athttps://github.com/wenjiegroup/iFORM.\n\nContactshuwj@bmi.ac.cn and boxc@bmi.ac.cn

Bioinformatics

A Structural and Functional View of Polypharmacology

Protein domains mediate drug-protein interactions and this effect can explain drug polypharmacology. In this study, we associate polypharmacological drugs with CATH functional families, a type of protein domain and we use the network properties of these druggable protein families to analyse their relationships with drug side effects. We found druggable CATH functional families enriched in drug targets, whose relatives are structurally coherent, gather together in the protein functional network occupying central positions, and tend to be free of proteins associated with drug side effects. Our results demonstrate that CATH functional families can be used to identify drug-target interactions, opening a new research direction in target identification.

Bioinformatics

NanoSim: nanopore sequence read simulator based on statistical characterization

MotivationIn 2014, Oxford Nanopore Technologies (ONT) announced a new sequencing platform called MinION. The particular features of MinION reads - longer read lengths and single-molecule sequencing in particular - show potential for genome characterization. As of yet, the pre-commercial technology is exclusively available through early-access, and only a few datasets are publically available for testing. Further, no software exists that simulates MinION platform reads with genuine ONT characteristics.\n\nResultsIn this article, we introduce NanoSim, a fast and scalable read simulator that captures the technology-specific features of ONT data, and allows for adjustments upon improvement of nanopore sequencing technology.

Bioinformatics

Automated SWATH Data Analysis Using Targeted Extraction of Ion Chromatograms

Targeted mass spectrometry comprises a set of methods able to quantify protein analytes in complex mixtures with high accuracy and sensitivity. These methods, e.g., Selected Reaction Monitoring (SRM) and SWATH MS, use specific mass spectrometric coordinates (assays) for reproducible detection and quantification of proteins. In this protocol, we describe how to analyze in a targeted manner data from a SWATH MS experiment aimed at monitoring thousands of proteins reproducibly over many samples. We present a standard SWATH MS analysis workflow, including manual data analysis for quality control (based on Skyline) as well as automated data analysis with appropriate control of error rates (based on the OpenSWATH workflow). We also discuss considerations to ensure maximal coverage, reproducibility and quantitative accuracy.

Bioinformatics

Zika virus outbreak in the Americas: Is Aedes albopictus an overlooked culprit?

Summary / AbstractCodon usage patterns of viruses reflect a series of evolutionary changes that enable viruses to shape their survival rates and fitness toward the external environment and, most importantly, their hosts. In the present study, we employed multiple codon usage analysis indices to determine genotype specific codon usage patterns of Zika virus (ZIKV) strains from the current outbreak and those reported previously. Several genotype specific and common codon usage traits were noted in ZIKV coding sequences, indicative of independent evolutionary origins from a common ancestor. The overall influence of natural selection was found to be more profound than that of mutation pressure and acting on specific set of viral genes belonging to ZIKV strains of Asian genotype from the recent outbreak. Furthermore, an interplay of codon adaptation and deoptimization have been observed in ZIKV genomes. The collective findings of codon analysis in association with the geographical data of Aedes populations in the Americas suggests that ZIKV have evolved a dynamic set of codon usage patterns in order to maintain a successful replication and transmission chain within multiple hosts and vectors.

Bioinformatics