Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,459 records · Page 81Linked to original sources

chromstaR: Tracking combinatorial chromatin state dynamics in space and time

BackgroundPost-translational modifications of histone residue tails are an important component of genome regulation. It is becoming increasingly clear that the combinatorial presence and absence of various modifications define discrete chromatin states which determine the functional properties of a locus. An emerging experimental goal is to track changes in chromatin state maps across different conditions, such as experimental treatments, cell-types or developmental time points.\n\nResultsHere we present chromstaR, an algorithm for the computational inference of combinatorial chromatin state dynamics across an arbitrary number of conditions. ChromstaR uses a multivariate Hidden Markov Model to determine the number of discrete combinatorial chromatin states using multiple ChIP-seq experiments as input and assigns every genomic region to a state based on the presence/absence of each modification in every condition. We demonstrate the advantages of chromstaR in the context of three common experimental data scenarios. First, we study how different histone modifications combine to form combinatorial chromatin states in a single tissue. Second, we infer genome-wide patterns of combinatorial state differences between two cell types or conditions. Finally, we study the dynamics of combinatorial chromatin states during tissue differentiation involving up to six differentiation points. Our findings reveal a striking sparcity in the combinatorial organization and temporal dynamics of chromatin state maps.\n\nConclusionschromstaR is a versatile computational tool that facilitates a deeper biological understanding of chromatin organization and dynamics. The algorithm is implemented as an R-package and freely available from http://bioconductor.org/packages/chromstaR/.

Bioinformatics

A novel approach to identifying marker genes and estimating the cellular composition of whole blood from gene expression profiles

Measuring genome-wide changes in transcript abundance in circulating peripheral whole blood cells is a useful way to study disease pathobiology and may help elucidate biomarkers and molecular mechanisms of disease. The sensitivity and interpretability of analyses carried out in this complex tissue, however, are significantly affected by its dynamic heterogeneity. It is therefore desirable to quantify this heterogeneity, either to account for it or to better model interactions that may be present between the abundance of certain transcripts, some cell types and the indication under study. Accurate enumeration of the many component cell types that make up peripheral whole blood can be costly, however, and may further complicate the sample collection process. Many approaches have been developed to infer the composition of a sample from high-dimensional transcriptomic and, more recently, epigenetic data. These approaches rely on the availability of isolated expression profiles for the cell types to be enumerated. These profiles are platform-specific, suitable datasets are rare, and generating them is expensive. No such dataset exists on the Affymetrix Gene ST platform. We present a freely-available, and open source, multi-response Gaussian model capable of accurately predicting the composition of peripheral whole blood samples from Affymetrix Gene ST expression profiles. This model outperforms other current methods when applied to Gene ST data and could potentially be used to enrich the >10,000 Affymetrix Gene ST blood gene expression profiles currently available on GEO.\n\nKey PointsO_LIWe introduce a model that accurately predicts the composition of blood from Affymetrix Gene ST gene expression profiles.\nC_LIO_LIThis model outperforms existing methods when applied to Affymetrix Gene ST expression profiles from blood.\nC_LI

Bioinformatics

sCNAphase: using haplotype resolved read depth to genotype somatic copy number alterations from low cellularity aneuploid tumors

Accurate identification of copy number alterations is an essential step in understanding the events driving tumor progression. While a variety of algorithms have been developed to use high-throughput sequencing data to profile copy number changes, no tool is able to reliably characterize ploidy and genotype absolute copy number from tumor samples which contain less than 40% tumor cells. To increase our power to resolve the copy number profile from low-cellularity tumor samples, we developed a novel approach which pre-phases heterozygote germline SNPs in order to replace the commonly used B-allele frequency with a more powerful parental-haplotype frequency. We apply our tool - sCNAphase - to characterize the copy number and loss-of-heterozygosity profiles of four publicly available breast cancer cell-lines. Comparisons to previous spectral karyotyping and microarray studies revealed that sCNAphase reliably identified overall ploidy as well as the individual copy number mutations from each cell-line. Analysis of artificial cell-line mixtures demonstrated the capacity of this method to determine the level of tumor cellularity, consistently identify sCNAs and characterize ploidy in samples with as little as 10% tumor cells. This novel methodology has the potential to bring sCNA profiling to low-cellularity tumors, a form of cancer unable to be accurately studied by current methods.

Bioinformatics

Partial derivatives meta-analysis: pooled analyses when individual participant data cannot be shared

Joint analysis of data from multiple studies in collaborative efforts strengthens scientific evidence, with the gold standard approach being the pooling of individual participant data (IPD). However, sharing IPD often has legal, ethical, and logistic constraints for sensitive or high-dimensional data, such as in clinical trials, observational studies, and large-scale omics studies. Therefore, meta-analysis of study-level effect estimates is routinely done, but this compromises on statistical power, accuracy, and flexibility. Here we propose a novel meta-analytical approach, named partial derivatives meta-analysis, that is mathematically equivalent to using IPD, yet only requires the sharing of aggregate data. It not only yields identical results as pooled IPD analyses, but also allows post-hoc adjustments for covariates and stratification without the need for site-specific re-analysis. Thus, in case that IPD cannot be shared, partial derivatives meta-analysis still produces gold standard results, which can be used to better inform guidelines and policies on clinical practice.

Bioinformatics

Multiple goal pursuit -- to kill two birds with one stone or to fall between two stools ?

We present the simple phenomenological (but - analytic) model allowing to formalize description of multitasking, i.e. simultaneous performing several tasks. That process requires distribution of attention, and for great number of goals do not lead to success. Our consideration shows that simultaneous performing more than two tasks is, most likely, impossible.

Bioinformatics

Implementation of an Open Source Software solution for Laboratory Information Management and automated RNAseq data analysis in a large-scale Cancer Genomics initiative using BASE with extension package Reggie.

BackgroundLarge-scale cancer genomics initiatives and next-generation sequencing for transcriptome profiling allow for detailed molecular characterization of tumors, and provide opportunities for clinical tools to improve diagnosis, prognosis, and treatment decisions. Laboratory information, data management, and data sharing in large-scale genomics projects is a challenge. Aiming to introduce such technologies in a clinical setting offer additional challenges associated with requirements of short lead-times and specialized tracking of biomaterials, data, and analysis results.\n\nResultsUsing the free open-source BioArray Software Environment (BASE) and extension package Reggie we have implemented a laboratory information management system and an automated RNAseq data analysis pipeline that successfully manage a large regional cancer genomics initiative. The system manages enrolled cancer patients, tumor biopsies, extraction of nucleic acid, and whole transcriptome RNA-sequencing through to data analysis and quality control. The implementation offers integration of laboratory equipment and operating procedures, and information tracking in a module based fashion enabling efficient and flexible use of personnel resources. The system provides two-factor authentication and transaction control and seamless integration of freely available software for RNAseq analysis such as Tophat, Cufflinks, and Picard. As of February 2016 more than 8000 patients and over 6000 tumor biopsies have been successfully processed. Lead-time from biopsy arrival to summarized reports based on RNAseq data is less than 5 days, in line with regional clinical requirements. BASE and Reggie are freely available and released as open-source under the GNU General Public License and GNU Affero General Public License, respectively.\n\nConclusionUsing free open-source software together with BASE and a customized extension package, Reggie, we have implemented a system capable of managing large collections of quality controlled and curated material for use in research and development and tailored to meet requirements for clinical use. Featuring high degree of automation and interactivity the system allows for resource efficient laboratory procedures and short lead-times with demonstrated use of RNAseq data analyses in a clinical setting.

Bioinformatics

Convert Your Favorite Protein Modeling Program Into A Mutation Predictor: "MODICT"

Motivation: Predict whether a mutation is deleterious based on the custom 3D model of a protein.\n\nMethods: We have developed O_SCPLOWMODIOTC_SCPLOW, a mutation prediction tool which is based on per residue RMSD (root mean square deviation) values of superimposed 3D protein models. Our mathematical algorithm was tested for 42 described mutations in multiple genes including renin, beta-tubulin, biotinidase, sphingomyelin phosphodiesterase-1, phenylalanine hydroxylase and medium chain Acyl-Coa dehydrogenase. Moreover, O_SCPLOWMODIOTC_SCPLOW scores corresponded to experimentally verified residual enzyme activities in mutated biotinidase, phenylalanine hydroxylase and medium chain Acyl-CoA dehydrogenase. Several commercially available prediction algorithms were tested and results were compared. The O_SCPLOWMODIOTC_SCPLOW PERL package and the manual can be downloaded from https://github.com/MODICT/MODICT.\n\nConclusion: We show here that O_SCPLOWMODIOTC_SCPLOW is capable tool for mutation effect prediction at the protein level, using superimposed 3D protein models instead of sequence based algorithms used by POLYPHEN and SIFT.

Bioinformatics

Bacterial sequences detected in 99 out of 99 serum samples from Ebola patients

Evolution and clinical manifestations of Ebola virus (EBOV) infection overlap with the pathologic processes that occur in sepsis1. Some viruses certainly compromise the immune system, leading to a breach in the integrity of the mucosal epithelial barrier, thus allowing bacterial translocation2, 3. Guided by these facts, we wondered if bacteria could be involved in the pathogenesis of some of the septic shock-like symptoms typical of EBOV infected patients, something that could have a dramatic impact on the design of new treatment approaches. We decided to search for bacteria in available EBOV patient sequence datasets. Given that EBOV is an RNA virus and that, hence, some NGS sequencing experiments carried out to sequence the EBOV genomes were RNA-Seq experiments, we thought that, if there were any bacteria in patient serum, at least some bacterial RNA might probably be detected in the sequenced material from Ebola patients. Thus, we searched for bacteria in a RNA-Seq public dataset from 99 Ebola samples from the last outbreak4, and surprisingly, in spite of the certainly suboptimal experimental conditions for bacterial RNA sequencing, we found bacteria in all of the 99 samples

Bioinformatics

IMP: a pipeline for reproducible metagenomic and metatranscriptomic analyses

We present IMP, an automated pipeline for reproducible integrated analyses of coupled metagenomic and metatranscriptomic data. IMP incorporates preprocessing, iterative co-assembly of metagenomic and metatranscriptomic data, analyses of microbial community structure and function as well as genomic signature-based visualizations. Complementary use of metagenomic and metatranscriptomic data improves assembly quality and enables the estimation of both population abundance and community activity while allowing the recovery and analysis of potentially important components, such as RNA viruses. IMP is containerized using Docker which ensures reproducibility. IMP is available at http://r3lab.uni.lu/web/imp/.

Bioinformatics

Detecting Horizontal Gene Transfer by Mapping Sequencing Reads Across Species Boundaries

Horizontal gene transfer (HGT) is a fundamental mechanism that enables organisms such as bacteria to directly transfer genetic material between distant species. This way, bacteria can acquire new traits such as antibiotic resistance or pathogenic toxins. Current bioinfor-matics approaches focus on the detection of past HGT events by exploring phylogenetic trees or genome composition inconsistencies. However, this normally requires the availability of finished and fully annotated genomes and of sufficiently large deviations that allow detection. Thus, these techniques are not widely applicable. Especially in an outbreak scenario where new HGT mediated pathogens emerge, there is need for fast and precise HGT detection. Next-generation sequencing (NGS) technologies can facilitate swift analysis of unknown pathogens but, to the best of our knowledge, so far no approach uses NGS data directly to detect HGTs.\n\nWe present Daisy, a novel mapping-based tool for HGT detection directly from NGS data. Daisy determines HGT boundaries with split-read mapping and evaluates candidate regions relying on read pair and coverage information. Daisy can successfully detect HGT regions with base pair resolution in both simulated and real data, and outperforms alternative approaches using a genome assembly of the reads. We see our approach as a powerful complement for a comprehensive analysis of HGT in the context of NGS data. Daisy is freely available from http://github.com/ktrappe/daisy.

Bioinformatics

Semi-Supervised Learning of the Electronic Health Record for Phenotype Stratification

Patient interactions with health care providers result in entries to electronic health records (EHRs). EHRs were built for clinical and billing purposes but contain many data points about an individual. Mining these records provides opportunities to extract electronic phenotypes, which can be paired with genetic data to identify genes underlying common human diseases. This task remains challenging: high quality phenotyping is costly and requires physician review; many fields in the records are sparsely filled; and our definitions of diseases are continuing to improve over time. Here we develop and evaluate a semi-supervised learning method for EHR phenotype extraction using denoising autoencoders for phenotype stratification. By combining denoising autoencoders with random forests we find classification improvements across multiple simulation models and improved survival prediction in ALS clinical trial data. This is particularly evident in cases where only a small number of patients have high quality phenotypes, a common scenario in EHR-based research. Denoising autoencoders perform dimensionality reduction enabling visualization and clustering for the discovery of new subtypes of disease. This method represents a promising approach to clarify disease subtypes and improve genotype-phenotype association studies that leverage EHRs.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC=\"FIGDIR/small/039800_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (31K):\norg.highwire.dtl.DTLVardef@567074org.highwire.dtl.DTLVardef@f0e45corg.highwire.dtl.DTLVardef@1207659org.highwire.dtl.DTLVardef@39f8b8_HPS_FORMAT_FIGEXP M_FIG C_FIG HIGHLIGHTSO_LIDenoising autoencoders (DAs) can model electronic health records.\nC_LIO_LISemi-supervised learning with DAs improves ALS patient survival predictions.\nC_LIO_LIDAs improve patient cluster visualization through dimensionality reduction.\nC_LI

Bioinformatics

MetaPalette: A K-mer painting approach for metagenomic taxonomic profiling and quantification of novel strain variation

Metagenomic profiling is challenging in part because of the highly uneven sampling of the tree of life by genome sequencing projects and the limitations imposed by performing phy-logenetic inference at fixed taxonomic ranks. We present the algorithm MetaPalette which uses long k-mer sizes (k = 30, 50) to fit a k-mer \"palette\" of a given sample to the k-mer palette of reference organisms. By modeling the k-mer palettes of unknown organisms, the method also gives an indication of the presence, abundance, and evolutionary relatedness of novel organisms present in the sample. The method returns a traditional, fixed-rank taxonomic profile which is shown on independently simulated data to be one of the most accurate to date. Tree figures are also returned that quantify the relatedness of novel organisms to reference sequences and the accuracy of such figures is demonstrated on simulated spike-ins and a metagenomic soil sample.\n\nThe software implementing MetaPalette is available at: https://github.com/dkoslicki/MetaPalette\n\nPre-trained databases are included for Archaea, Bacteria, Eukaryota, and viruses.

Bioinformatics

Succinct Colored de Bruijn Graphs

Iqbal et al. (Nature Genetics, 2012) introduced the colored de Bruijn graph, a variant of the classic de Bruijn graph, which is aimed at \"detecting and genotyping simple and complex genetic variants in an individual or population\". Because they are intended to be applied to massive population level data, it is essential that the graphs be represented efficiently. Unfortunately, current succinct de Bruijn graph representations are not directly applicable to the colored de Bruijn graph, which require additional information to be succinctly encoded as well as support for non-standard traversal operations. Our data structure dramatically reduces the amount of memory required to store and use the colored de Bruijn graph, with some penalty to runtime, allowing it to be applied in much larger and more ambitious sequence projects than was previously possible.

Bioinformatics

Combating Chagas Disease Through Inhibition of Tiam1, a Rho GTPase Guanine Nucleotide Exchange Factor

Chagas disease is a major cardiovascular affliction primarily endemic to Latin American countries, affecting some ten to twelve million people worldwide. The currently available drugs, Benznidazole and Nifurtimox, are ineffective in the chronic stages and induce severe side effects. In an attempt to improve this situation we use an in silico drug repurposing strategy to correlate drug-protein interactions with positive clinical outcomes. The strategy involves a protein functional site similarity search, along with computational docking studies and, given the findings, a phosphatidylinositol (PIP) strip test to determine the activity of Posaconazole, a recently developed antifungal triazole, in conjunction with Tiam1, a Rho GTPase Guanine Nucleotide Exchange Factor. The results from both computational and in vitro studies indicate possible inhibition of phosphoinositides via Posaconazole, preventing Rho GTPase-induced proliferation of T. cruzi, the etiological agent of Chagas Disease.

Bioinformatics

variancePartition: Interpreting drivers of variation in complex gene expression studies

As genomics studies become more complex and consider multiple sources of biological and technical variation, characterizing these drivers of variation becomes essential to understanding disease biology and regulatory genetics. We describe a statistical and visualization framework, variancePartition, to prioritize drivers of variation with a genome-wide summary, and identify genes that deviate from the genome-wide trend. variancePartition enables rapid interpretation of complex gene expression studies and is applicable to many genomics assays.

Bioinformatics

Reconstruction of ancestral genomes in presence of gene gain and loss.

Since most dramatic genomic changes are caused by genome rearrangements as well as gene duplications and gain/loss events, it becomes crucial to understand their mechanisms and reconstruct ancestral genomes of the given genomes. This problem was shown to be NP-complete even in the \"simplest\" case of three genomes, thus calling for heuristic rather than exact algorithmic solutions. At the same time, a larger number of input genomes may actually simplify the problem in practice as it was earlier illustrated with MGRA, a state-of-the-art software tool for reconstruction of ancestral genomes of multiple genomes.\n\nOne of the key obstacles for MGRA and other similar tools is presence of breakpoint reuses when the same breakpoint region is broken by several different genome rearrangements in the course of evolution. Furthermore, such tools are often limited to genomes composed of the same genes with each gene present in a single copy in every genome. This limitation makes these tools inapplicable for many biological datasets and degrades the resolution of ancestral reconstructions in diverse datasets.\n\nWe address these deficiencies by extending the MGRA algorithm to genomes with unequal gene contents. The developed next-generation tool MGRA2 can handle gene gain/loss events and shares the ability of MGRA to reconstruct ancestral genomes uniquely in the case of limited breakpoint reuse. Furthermore, MGRA2 employs a number of novel heuristics to cope with higher breakpoint reuse and process datasets inaccessible for MGRA. In practical experiments, MGRA2 shows superior performance for simulated and real genomes as compared to other ancestral genomes reconstruction tools. The MGRA2 tool is distributed as an open-source software and can be downloaded from GitHub repository http://github.com/ablab/mgra/. It is also available in the form of a web-server at http://mgra.cblab.org, which makes it readily accessible for inexperienced users.

Bioinformatics

StrainSeeker: fast identification of bacterial strains from unassembled sequencing reads using user-provided guide trees.

BackgroundFast, accurate and high-throughput detection of bacteria is in great demand. The present work was conducted to investigate the possibility of identifying both known and unknown bacterial strains from unassembled next-generation sequencing reads using custom-made guide trees.\n\nResultsA program named StrainSeeker was developed that constructs a list of specific k-mers for each node of any given Newick-format tree and enables rapid identification of bacterial genomes within minutes. StrainSeeker has been tested and shown to successfully identify Escherichia coli strains from mixed samples in less than 5 minutes. StrainSeeker can also identify bacterial strains from highly diverse metagenomics samples. StrainSeeker is available at http://bioinfo.ut.ee/strainseeker.\n\nConclusionsOur novel approach can be useful for both clinical diagnostics and research laboratories because novel bacterial strains are constantly emerging and their fast and accurate detection is very important.

Bioinformatics

AC-PCA: simultaneous dimension reduction and adjustment for confounding variation

Dimension reduction methods are commonly applied to high-throughput biological datasets. However, the results can be hindered by confounding factors, either biologically or technically originated. In this study, we extend Principal Component Analysis to propose AC-PCA for simultaneous dimension reduction and adjustment for confounding variation. We show that AC-PCA can adjust for a) variations across individual donors present in a human brain exon array dataset, and b) variations of different species in a model organism ENCODE RNA-Seq dataset. Our approach is able to recover the anatomical structure of neocortical regions, and to capture the shared variation among species during embryonic development. For gene selection purposes, we extend AC-PCA with sparsity constraints, and propose and implement an efficient algorithm. The methods developed in this paper can also be applied to more general settings.

Bioinformatics