Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76Linked to original sources

SNP-sites: rapid efficient extraction of SNPs from multi-FASTA alignments

Rapidly decreasing genome sequencing costs have led to a proportionate increase in the number of samples used in prokaryotic population studies. Extracting single nucleotide polymorphisms (SNPs) from a large whole genome alignment is now a routine task, but existing tools have failed to scale efficiently with the increased size of studies. These tools are slow, memory inefficient and are installed through non-standard procedures. We present SNP-sites which can rapidly extract SNPs from a multi-FASTA alignment using modest resources and can output results in multiple formats for downstream analysis. SNPs can be extracted from a 8.3 GB alignment file (1,842 taxa, 22,618 sites) in 267 seconds using 59 MB of RAM and 1 CPU core, making it feasible to run on modest computers. It is easy to install through the Debian and Homebrew package managers, and has been successfully tested on more than 20 operating systems. SNP-sites is implemented in C and is available under the open source license GNU GPL version 3.

Bioinformatics

MetaFlow: Metagenomic profiling based on whole-genome coverage analysis with min-cost flows

High-throughput sequencing (HTS) of metagenomes is proving essential in understanding the environment and diseases. State-of-the-art methods for discovering the species and their abundances in an HTS metagenomic sample are based on genome-specific markers, which can lead to skewed results, especially at species level. We present MetaFlow, the first method based on coverage analysis across entire genomes that also scales to HTS samples. We formulated this problem as an NP-hard matching problem in a bipartite graph, which we solved in practice by min-cost flows. On synthetic data sets of varying complexity and similarity, MetaFlow is more precise and sensitive than popular tools such as MetaPhlAn, mOTU, GSMer and BLAST, and its abundance estimations at species level are two to four times better in terms of{ell} 1-norm. On a real human stool data set, MetaFlow identifies B.uniformis as most predominant, in line with previous human gut studies, whereas marker-based methods report it as rare. MetaFlow is freely available at http://cs.helsinki.fi/gsa/metaflow

Bioinformatics

False Discovery Rates: A New Deal.

We introduce a new Empirical Bayes approach for large-scale hypothesis testing, including estimating False Discovery Rates (FDRs), and effect sizes. This approach has two key differences from existing approaches to FDR analysis. First, it assumes that the distribution of the actual (unobserved) effects is unimodal, with a mode at 0. This \"unimodal assumption\" (UA), although natural in many contexts, is not usually incorporated into standard FDR analysis, and we demonstrate how incorporating it brings many benefits. Specifically, the UA facilitates efficient and robust computation - estimating the unimodal distribution involves solving a simple convex optimization problem - and enables more accurate inferences provided that it holds. Second, the method takes as its input two numbers for each test (an effect size estimate, and corresponding standard error), rather than the one number usually used (p value, or z score). When available, using two numbers instead of one helps account for variation in measurement precision across tests. It also facilitates estimation of effects, and unlike standard FDR methods our approach provides interval estimates (credible regions) for each effect in addition to measures of significance. To provide a bridge between interval estimates and significance measures we introduce the term \"local false sign rate\" to refer to the probability of getting the sign of an effect wrong, and argue that it is a superior measure of significance than the local FDR because it is both more generally applicable, and can be more robustly estimated. Our methods are implemented in an R package ashr available from http://github.com/stephens999/ashr.

Bioinformatics

Application of new informatics tools for identifying allosteric lead ligands of the c-Src kinase

Recent molecular dynamics (MD) simulations of the catalytic domain of the c-Src kinase revealed intermediate conformations with a potentially druggable allosteric pocket adjacent to the C-helix, bound by 8-anilino-1-naphthalene sulfonate. Towards confirming the existence of this pocket, we have developed a novel lead enrichment protocol using new target and lead enrichment software to identify sixteen allosteric lead ligands of the c-Src kinase. First, Markov State Models analysis was used to identify the most statistically significant c-Src target conformations from all MD-simulated conformations. The most statistically relevant candidate MSM targets were then prioritized by assessing how well each reproduced binding poses of ligands specific to the ATP-competitive and allosteric pockets. The top-performing MSM targets, identified by receiver-operating curve analysis, were then used to screen the ZINC library of 13 million clean, drug-like ligands, all of which prioritized based on their empirical scoring function, binding pose consistency across MSM targets, and strong hydrogen bonding and hydrophobic interactions with Src residues. The FragFEATURE knowledgebase of fragment-protein pocket interactions was then used to identify fragments specific to the ATP-competitive and allosteric pockets. This information was used to identify seven Type II and nine Type III lead ligands with binding poses supported by fragment predictions. Of these, Type II lead ligands, ZINC13037947 and ZINC09672647, and Type III lead ligands, ZINC12530852 and ZINC30012975, exhibited the most favorable fragment profiles and are recommended for further experimental testing for the existence of the allosteric pocket in Src.

Bioinformatics

msVolcano: a flexible web application for visualizing quantitative proteomics data.

We introduce msVolcano, a web application, for the visualization of label-free mass spectrometric data. It is optimized for the output of the MaxQuant data analysis pipeline of interactomics experiments and generates volcano plots with lists of interacting proteins. The user can optimize the cutoff values to find meaningful significant interactors for the tagged protein of interest. Optionally, stoichiometries of interacting proteins can be calculated. Several customization options are provided to the user for flexibility and publication-quality outputs can also be downloaded (tabular and graphical).\n\nAvailability: msVolcano is implemented in R Statistical language using Shiny and is hosted at server in-house. It can be accessed freely from anywhere at http://projects.biotec.tudresden.de/msVolcano/

Bioinformatics

Predicting protein thermal stability changes upon point mutations using statistical potentials: Introducing HoTMuSiC

The accurate prediction of the impact of an amino acid substitution on the thermal stability of a protein is a central issue in protein science, and is of key relevance for the rational optimization of various bioprocesses that use enzymes in unusual conditions. Here we present one of the first computational tools to predict the change in melting temperature {Delta}Tm upon point mutations, given the protein structure and, when available, the melting temperature Tm of the wild-type protein. The key ingredients of our model structure are standard and temperature-dependent statistical potentials, which are combined with the help of an artificial neural network. The model structure was chosen on the basis of a detailed thermodynamic analysis of the system. The parameters of the model were identified on a set of more than 1,600 mutations with experimentally measured {Delta}Tm. The performance of our method was tested using a strict 5-fold cross-validation procedure, and was found to be significantly superior to that of competing methods. We obtained a root mean square deviation between predicted and experimental {Delta}Tm values of 4.2{degrees}C that reduces to 2.9{degrees}C when ten percent outliers are removed. A webserver-based tool is freely available for non-commercial use at soft.dezyme.com.

Bioinformatics

Information-dependent Enrichment Analysis Reveals Time-dependent Transcriptional Regulation of the Estrogen Pathway of Toxicity

The twenty-first century vision for toxicology involves a transition away from high-dose animal studies and into in vitro and computational models. This movement requires mapping pathways of toxicity through an understanding of how in vitro systems respond to chemical perturbation. Uncovering transcription factors responsible for gene expression patterns is essential for defining pathways of toxicity, and ultimately, for determining chemical mode of action, through which a toxicant acts. Traditionally this is achieved via chromatin immunoprecipitation studies and summarized by calculating, which transcription factors are statistically associated with the up-and down-regulated genes. These lists are commonly determined via statistical or fold-change cutoffs, a procedure that is sensitive to statistical power and may not be relevant to determining transcription factor associations. To move away from an arbitrary statistical or fold-change based cutoffs, we have developed in the context of the Mapping the Human Toxome project, a novel enrichment paradigm called Information Dependent Enrichment Analysis (IDEA) to guide identification of the transcription factor network. We used the test case of endocrine disruption of MCF-7 cells activated by 17{beta} estradiol (E2). Using this new approach, we were able to establish a time course for transcriptional and functional responses to E2. ER and ER{beta} are associated with short-term transcriptional changes in response to E2. Sustained exposure leads to the recruitment of an additional ensemble of transcription factors and alteration of cell-cycle machinery. TFAP2C and SOX2 were the transcription factors most highly correlated with dose. E2F7, E2F1 and Foxm1, which are involved in cell proliferation, were enriched only at 24h. IDEA is, therefore, a novel tool to identify candidate pathways of toxicity, clearly outperforming Gene-set Enrichment Analysis but with similar results as Weighted Gene Correlation Network Analysis, which helps to identify genes not annotated to pathways.

Bioinformatics

chromstaR: Tracking combinatorial chromatin state dynamics in space and time

BackgroundPost-translational modifications of histone residue tails are an important component of genome regulation. It is becoming increasingly clear that the combinatorial presence and absence of various modifications define discrete chromatin states which determine the functional properties of a locus. An emerging experimental goal is to track changes in chromatin state maps across different conditions, such as experimental treatments, cell-types or developmental time points.\n\nResultsHere we present chromstaR, an algorithm for the computational inference of combinatorial chromatin state dynamics across an arbitrary number of conditions. ChromstaR uses a multivariate Hidden Markov Model to determine the number of discrete combinatorial chromatin states using multiple ChIP-seq experiments as input and assigns every genomic region to a state based on the presence/absence of each modification in every condition. We demonstrate the advantages of chromstaR in the context of three common experimental data scenarios. First, we study how different histone modifications combine to form combinatorial chromatin states in a single tissue. Second, we infer genome-wide patterns of combinatorial state differences between two cell types or conditions. Finally, we study the dynamics of combinatorial chromatin states during tissue differentiation involving up to six differentiation points. Our findings reveal a striking sparcity in the combinatorial organization and temporal dynamics of chromatin state maps.\n\nConclusionschromstaR is a versatile computational tool that facilitates a deeper biological understanding of chromatin organization and dynamics. The algorithm is implemented as an R-package and freely available from http://bioconductor.org/packages/chromstaR/.

Bioinformatics

A novel approach to identifying marker genes and estimating the cellular composition of whole blood from gene expression profiles

Measuring genome-wide changes in transcript abundance in circulating peripheral whole blood cells is a useful way to study disease pathobiology and may help elucidate biomarkers and molecular mechanisms of disease. The sensitivity and interpretability of analyses carried out in this complex tissue, however, are significantly affected by its dynamic heterogeneity. It is therefore desirable to quantify this heterogeneity, either to account for it or to better model interactions that may be present between the abundance of certain transcripts, some cell types and the indication under study. Accurate enumeration of the many component cell types that make up peripheral whole blood can be costly, however, and may further complicate the sample collection process. Many approaches have been developed to infer the composition of a sample from high-dimensional transcriptomic and, more recently, epigenetic data. These approaches rely on the availability of isolated expression profiles for the cell types to be enumerated. These profiles are platform-specific, suitable datasets are rare, and generating them is expensive. No such dataset exists on the Affymetrix Gene ST platform. We present a freely-available, and open source, multi-response Gaussian model capable of accurately predicting the composition of peripheral whole blood samples from Affymetrix Gene ST expression profiles. This model outperforms other current methods when applied to Gene ST data and could potentially be used to enrich the >10,000 Affymetrix Gene ST blood gene expression profiles currently available on GEO.\n\nKey PointsO_LIWe introduce a model that accurately predicts the composition of blood from Affymetrix Gene ST gene expression profiles.\nC_LIO_LIThis model outperforms existing methods when applied to Affymetrix Gene ST expression profiles from blood.\nC_LI

Bioinformatics

sCNAphase: using haplotype resolved read depth to genotype somatic copy number alterations from low cellularity aneuploid tumors

Accurate identification of copy number alterations is an essential step in understanding the events driving tumor progression. While a variety of algorithms have been developed to use high-throughput sequencing data to profile copy number changes, no tool is able to reliably characterize ploidy and genotype absolute copy number from tumor samples which contain less than 40% tumor cells. To increase our power to resolve the copy number profile from low-cellularity tumor samples, we developed a novel approach which pre-phases heterozygote germline SNPs in order to replace the commonly used B-allele frequency with a more powerful parental-haplotype frequency. We apply our tool - sCNAphase - to characterize the copy number and loss-of-heterozygosity profiles of four publicly available breast cancer cell-lines. Comparisons to previous spectral karyotyping and microarray studies revealed that sCNAphase reliably identified overall ploidy as well as the individual copy number mutations from each cell-line. Analysis of artificial cell-line mixtures demonstrated the capacity of this method to determine the level of tumor cellularity, consistently identify sCNAs and characterize ploidy in samples with as little as 10% tumor cells. This novel methodology has the potential to bring sCNA profiling to low-cellularity tumors, a form of cancer unable to be accurately studied by current methods.

Bioinformatics

Partial derivatives meta-analysis: pooled analyses when individual participant data cannot be shared

Joint analysis of data from multiple studies in collaborative efforts strengthens scientific evidence, with the gold standard approach being the pooling of individual participant data (IPD). However, sharing IPD often has legal, ethical, and logistic constraints for sensitive or high-dimensional data, such as in clinical trials, observational studies, and large-scale omics studies. Therefore, meta-analysis of study-level effect estimates is routinely done, but this compromises on statistical power, accuracy, and flexibility. Here we propose a novel meta-analytical approach, named partial derivatives meta-analysis, that is mathematically equivalent to using IPD, yet only requires the sharing of aggregate data. It not only yields identical results as pooled IPD analyses, but also allows post-hoc adjustments for covariates and stratification without the need for site-specific re-analysis. Thus, in case that IPD cannot be shared, partial derivatives meta-analysis still produces gold standard results, which can be used to better inform guidelines and policies on clinical practice.

Bioinformatics

Multiple goal pursuit -- to kill two birds with one stone or to fall between two stools ?

We present the simple phenomenological (but - analytic) model allowing to formalize description of multitasking, i.e. simultaneous performing several tasks. That process requires distribution of attention, and for great number of goals do not lead to success. Our consideration shows that simultaneous performing more than two tasks is, most likely, impossible.

Bioinformatics

Implementation of an Open Source Software solution for Laboratory Information Management and automated RNAseq data analysis in a large-scale Cancer Genomics initiative using BASE with extension package Reggie.

BackgroundLarge-scale cancer genomics initiatives and next-generation sequencing for transcriptome profiling allow for detailed molecular characterization of tumors, and provide opportunities for clinical tools to improve diagnosis, prognosis, and treatment decisions. Laboratory information, data management, and data sharing in large-scale genomics projects is a challenge. Aiming to introduce such technologies in a clinical setting offer additional challenges associated with requirements of short lead-times and specialized tracking of biomaterials, data, and analysis results.\n\nResultsUsing the free open-source BioArray Software Environment (BASE) and extension package Reggie we have implemented a laboratory information management system and an automated RNAseq data analysis pipeline that successfully manage a large regional cancer genomics initiative. The system manages enrolled cancer patients, tumor biopsies, extraction of nucleic acid, and whole transcriptome RNA-sequencing through to data analysis and quality control. The implementation offers integration of laboratory equipment and operating procedures, and information tracking in a module based fashion enabling efficient and flexible use of personnel resources. The system provides two-factor authentication and transaction control and seamless integration of freely available software for RNAseq analysis such as Tophat, Cufflinks, and Picard. As of February 2016 more than 8000 patients and over 6000 tumor biopsies have been successfully processed. Lead-time from biopsy arrival to summarized reports based on RNAseq data is less than 5 days, in line with regional clinical requirements. BASE and Reggie are freely available and released as open-source under the GNU General Public License and GNU Affero General Public License, respectively.\n\nConclusionUsing free open-source software together with BASE and a customized extension package, Reggie, we have implemented a system capable of managing large collections of quality controlled and curated material for use in research and development and tailored to meet requirements for clinical use. Featuring high degree of automation and interactivity the system allows for resource efficient laboratory procedures and short lead-times with demonstrated use of RNAseq data analyses in a clinical setting.

Bioinformatics

Convert Your Favorite Protein Modeling Program Into A Mutation Predictor: "MODICT"

Motivation: Predict whether a mutation is deleterious based on the custom 3D model of a protein.\n\nMethods: We have developed O_SCPLOWMODIOTC_SCPLOW, a mutation prediction tool which is based on per residue RMSD (root mean square deviation) values of superimposed 3D protein models. Our mathematical algorithm was tested for 42 described mutations in multiple genes including renin, beta-tubulin, biotinidase, sphingomyelin phosphodiesterase-1, phenylalanine hydroxylase and medium chain Acyl-Coa dehydrogenase. Moreover, O_SCPLOWMODIOTC_SCPLOW scores corresponded to experimentally verified residual enzyme activities in mutated biotinidase, phenylalanine hydroxylase and medium chain Acyl-CoA dehydrogenase. Several commercially available prediction algorithms were tested and results were compared. The O_SCPLOWMODIOTC_SCPLOW PERL package and the manual can be downloaded from https://github.com/MODICT/MODICT.\n\nConclusion: We show here that O_SCPLOWMODIOTC_SCPLOW is capable tool for mutation effect prediction at the protein level, using superimposed 3D protein models instead of sequence based algorithms used by POLYPHEN and SIFT.

Bioinformatics

Bacterial sequences detected in 99 out of 99 serum samples from Ebola patients

Evolution and clinical manifestations of Ebola virus (EBOV) infection overlap with the pathologic processes that occur in sepsis1. Some viruses certainly compromise the immune system, leading to a breach in the integrity of the mucosal epithelial barrier, thus allowing bacterial translocation2, 3. Guided by these facts, we wondered if bacteria could be involved in the pathogenesis of some of the septic shock-like symptoms typical of EBOV infected patients, something that could have a dramatic impact on the design of new treatment approaches. We decided to search for bacteria in available EBOV patient sequence datasets. Given that EBOV is an RNA virus and that, hence, some NGS sequencing experiments carried out to sequence the EBOV genomes were RNA-Seq experiments, we thought that, if there were any bacteria in patient serum, at least some bacterial RNA might probably be detected in the sequenced material from Ebola patients. Thus, we searched for bacteria in a RNA-Seq public dataset from 99 Ebola samples from the last outbreak4, and surprisingly, in spite of the certainly suboptimal experimental conditions for bacterial RNA sequencing, we found bacteria in all of the 99 samples

Bioinformatics

IMP: a pipeline for reproducible metagenomic and metatranscriptomic analyses

We present IMP, an automated pipeline for reproducible integrated analyses of coupled metagenomic and metatranscriptomic data. IMP incorporates preprocessing, iterative co-assembly of metagenomic and metatranscriptomic data, analyses of microbial community structure and function as well as genomic signature-based visualizations. Complementary use of metagenomic and metatranscriptomic data improves assembly quality and enables the estimation of both population abundance and community activity while allowing the recovery and analysis of potentially important components, such as RNA viruses. IMP is containerized using Docker which ensures reproducibility. IMP is available at http://r3lab.uni.lu/web/imp/.

Bioinformatics

Detecting Horizontal Gene Transfer by Mapping Sequencing Reads Across Species Boundaries

Horizontal gene transfer (HGT) is a fundamental mechanism that enables organisms such as bacteria to directly transfer genetic material between distant species. This way, bacteria can acquire new traits such as antibiotic resistance or pathogenic toxins. Current bioinfor-matics approaches focus on the detection of past HGT events by exploring phylogenetic trees or genome composition inconsistencies. However, this normally requires the availability of finished and fully annotated genomes and of sufficiently large deviations that allow detection. Thus, these techniques are not widely applicable. Especially in an outbreak scenario where new HGT mediated pathogens emerge, there is need for fast and precise HGT detection. Next-generation sequencing (NGS) technologies can facilitate swift analysis of unknown pathogens but, to the best of our knowledge, so far no approach uses NGS data directly to detect HGTs.\n\nWe present Daisy, a novel mapping-based tool for HGT detection directly from NGS data. Daisy determines HGT boundaries with split-read mapping and evaluates candidate regions relying on read pair and coverage information. Daisy can successfully detect HGT regions with base pair resolution in both simulated and real data, and outperforms alternative approaches using a genome assembly of the reads. We see our approach as a powerful complement for a comprehensive analysis of HGT in the context of NGS data. Daisy is freely available from http://github.com/ktrappe/daisy.

Bioinformatics

Semi-Supervised Learning of the Electronic Health Record for Phenotype Stratification

Patient interactions with health care providers result in entries to electronic health records (EHRs). EHRs were built for clinical and billing purposes but contain many data points about an individual. Mining these records provides opportunities to extract electronic phenotypes, which can be paired with genetic data to identify genes underlying common human diseases. This task remains challenging: high quality phenotyping is costly and requires physician review; many fields in the records are sparsely filled; and our definitions of diseases are continuing to improve over time. Here we develop and evaluate a semi-supervised learning method for EHR phenotype extraction using denoising autoencoders for phenotype stratification. By combining denoising autoencoders with random forests we find classification improvements across multiple simulation models and improved survival prediction in ALS clinical trial data. This is particularly evident in cases where only a small number of patients have high quality phenotypes, a common scenario in EHR-based research. Denoising autoencoders perform dimensionality reduction enabling visualization and clustering for the discovery of new subtypes of disease. This method represents a promising approach to clarify disease subtypes and improve genotype-phenotype association studies that leverage EHRs.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC=\"FIGDIR/small/039800_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (31K):\norg.highwire.dtl.DTLVardef@567074org.highwire.dtl.DTLVardef@f0e45corg.highwire.dtl.DTLVardef@1207659org.highwire.dtl.DTLVardef@39f8b8_HPS_FORMAT_FIGEXP M_FIG C_FIG HIGHLIGHTSO_LIDenoising autoencoders (DAs) can model electronic health records.\nC_LIO_LISemi-supervised learning with DAs improves ALS patient survival predictions.\nC_LIO_LIDAs improve patient cluster visualization through dimensionality reduction.\nC_LI

Bioinformatics