Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Accurate informatic modeling of tooth enamel pellicle interactions by training substitution matrices with Mat4Pep

MotivationProtein-hydroxyapatite interactions govern the development and homeostasis of teeth and bone. Characterization would enable design of peptides to regenerate mineralized tissues and control attachments such as ligaments and dental plaque. Progress has been limited because no available methods produce robust data for assessing phase interfaces.\n\nResultsWe show that tooth enamel pellicle peptides contain subtle sequence similarities that encode hydroxyapatite binding mechanisms, by segregating pellicle peptides from control sequences using our previously developed substitution matrix-based peptide comparison protocol (Oren et al., 2007), with improvements. Sampling diverse matrices, adding biological control sequences, and optimizing matrix refinement algorithms improves discrimination from 0.81 to 0.99 AUC in leave-one-out experiments. Other contemporary methods fail on this problem. We find hydroxyapatite interaction sequence patterns by applying the resulting selected refined matrix (\"pellitrix\") to cluster the peptides and build subgroup alignments. We identify putative hydroxyapatite maturation domains by application to enamel biomineralization proteins and prioritize putative novel pellicle peptides identified by In stageTip (iST) mass spectrometry. The sequence comparison protocol outperforms other contemporary options for this small and heterogeneous group, and is generalized for application to any group of peptides.\n\nAvailabilitySoftware to apply this protocol is freely available at github.com/JeremyHorst/Mat4Pep and compbio.org/protinfo/ Mat4Pep.\n\nContactjahorst@gmail.com, ram@compbio.org.\n\nSupplementary informationAvailable at Bioinformatics online.

bioinformatics

Critical assessment of approaches for molecular docking to elucidate associations of HLA alleles with Adverse Drug Reactions

Adverse drug reactions have been linked with genetic polymorphisms in HLA genes in numerous different studies. HLA proteins have an essential role in the presentation of self and non-self peptides, as part of the adaptive immune response. Amongst the associated drugs-allele combinations, anti-HIV drug Abacavir has been shown to be associated with the HLA-B*57:01 allele, and anti-epilepsy drug Carbamazepine with B*15:02, in both cases likely following the altered peptide repertoire model of interaction. Under this model, the drug binds directly to the antigen presentation region, causing different self peptides to be presented, which trigger an unwanted immune response. There is growing interest in searching for evidence supporting this model for other ADRs using bioinformatics techniques. In this study, in silico docking was used to assess the utility and reliability of well-known docking programs when addressing these challenging HLA-drug situations. Four docking programs: SwissDock, ROSIE, AutoDock Vina and AutoDockFR, were used to investigate if each software could accurately dock the Abacavir back into the crystal structure for the protein arising from the known risk allele, and if they were able to distinguish between the HLA-associated and non-HLA-associated (control) alleles. The impact of using homology models on the docking performance and how using different parameters such as including receptor flexibility affected the docking performance, were also investigated to simulate the approach where a crystal structure for a given HLA allele may be unavailable. The programs that were best able to predict the binding position of Abacavir were then used to recreate the docking seen for Carbamazepine with B*15:02 and controls alleles. It was found that the programmes investigated were sometimes able to correctly predict the binding mode of Abacavir with B*57:01 but not always. Each of the software packages that were assessed could predict the binding of Abacavir and Carbamazepine within the correct sub-pocket and, with the exception of ROSIE, was able to correctly distinguish between risk and control alleles. We found that docking to homology models could produce poorer quality predictions, especially when sequence differences impact the architecture of predicted binding pockets. Caution must therefore be used as inaccurate structures may lead to erroneous docking predictions. Incorporating receptor flexibility was found to negatively affect the docking performance for the examples investigated. Taken together, our findings help characterise the potential but also the limitations of computational prediction of drug-HLA interactions. These docking techniques should therefore always be used with care and alongside other methods of investigation, in order to be able to draw strong conclusions from the given results.

bioinformatics

Methods for Automatic Reference Trees and Multilevel Phylogenetic Placement

MotivationIn most metagenomic sequencing studies, the initial analysis step consists in assessing the evolutionary provenance of the sequences. Phylogenetic (or Evolutionary) Placement methods can be employed to determine the evolutionary position of sequences with respect to a given reference phylogeny. These placement methods do however face certain limitations: The manual selection of reference sequences is labor-intensive; the computational effort to infer reference phylogenies is substantially larger than for methods that rely on sequence similarity; the number of taxa in the reference phylogeny should be small enough to allow for visually inspecting the results.\n\nResultsWe present algorithms to overcome the above limitations. First, we introduce a method to automatically construct representative sequences from databases to infer reference phylogenies. Second, we present an approach for conducting large-scale phylogenetic placements on nested phylogenies. Third, we describe a preprocessing pipeline that allows for handling huge sequence data sets. Our experiments on empirical data show that our methods substantially accelerate the workflow and yield highly accurate placement results.\n\nImplementationFreely available under GPLv3 at http://github.com/lczech/gappa.\n\nContactlucas.czech@h-its.org\n\nSupplementary InformationSupplementary data are available at Bioinformatics online.

bioinformatics

Optimum Search Schemes for Approximate String Matching Using Bidirectional FM-Index

Finding approximate occurrences of a pattern in a text using a full-text index is a central problem in bioinformatics and has been extensively researched. Bidirectional indices have opened new possibilities in this regard allowing the search to start from anywhere within the pattern and extend in both directions. In particular, use of search schemes (partitioning the pattern and searching the pieces in certain orders with given bounds on errors) can yield significant speed-ups. However, finding optimal search schemes is a difficult combinatorial optimization problem.\n\nHere for the first time, we propose a mixed integer program (MIP) capable to solve this optimization problem for Hamming distance with given number of pieces. Our experiments show that the optimal search schemes found by our MIP significantly improve the performance of search in bidirectional FM-index upon previous ad-hoc solutions. For example, approximate matching of 101-bp Illumina reads (with two errors) becomes 35 times faster than standard backtracking. Moreover, despite being performed purely in the index, the running time of search using our optimal schemes (for up to two errors) is comparable to the best state-of-the-art aligners, which benefit from combining search in index with in-text verification using dynamic programming. As a result, we anticipate a full-fledged aligner that employs an intelligent combination of search in the bidirectional FM-index using our optimal search schemes and in-text verification using dynamic programming that will outperform todays best aligners. The development of such an aligner, called FAMOUS (Fast Approximate string Matching using OptimUm search Schemes), is ongoing as our future work.

bioinformatics

MGDb: An analyzed database and a genomic resource of mango (Mangifera Indica L.) cultivars for mango research

Mango is one of the famous and fifth most important subtropical/tropical fruit crops worldwide with the production centered in India and South-East Asia. Recently, there has been a worldwide interest in mango genomics to produce tools for Marker Assisted Selection and trait association. There are no web-based analyzed genomic resources available for mango particularly. Hence a complete mango genomic resource was required for improvement in research and management of mango germplasm. In this project, we have done comparative transcriptome analysis of four mango cultivars i.e. cv. Langra, cv. Zill, cv. Shelly and cv. Kent from Pakistan, China, Israel, and Mexico respectively. The raw data is obtained through De-novo sequence assembly which generated 30,953-85,036 unigenes from RNA-Seq datasets of mango cultivars. The project is aimed to provide the scientific community and general public a mango genomic resource and allow the user to examine their data against our analyzed mango genome databases of four cultivars (cv. Langra, cv. Zill, cv. Shelly and cv. Kent). A mango web genomic resource MGdb, is based on 3-tier architecture, developed using Python, flat file database, and JavaScript. It contains the information of predicted genes of the whole genome, the unigenes annotated by homologous genes in other species, and GO (Gene Ontology) terms which provide a glimpse of the traits in which they are involved. This web genomic resource can be of immense use in the assessment of the research, development of the medicines, understanding genetics and provides useful bioinformatics solution for analysis of nucleotide sequence data. We report here worlds first web-based genomic resource particularly of mango for genetic improvement and management of mango genome.

bioinformatics

emeraLD: Rapid Linkage Disequilibrium Estimation with Massive Data Sets

SummaryEstimating linkage disequilibrium (LD) is essential for a wide range of summary statistics-based association methods for genome-wide association studies (GWAS). Large genetic data sets, e.g. the TOPMed WGS project and UK Biobank, enable more accurate and comprehensive LD estimates, but increase the computational burden of LD estimation. Here, we describe emeraLD (Efficient Methods for Estimation and Random Access of LD), a computational tool that leverages sparsity and haplotype structure to estimate LD orders of magnitude faster than existing tools.\n\nAvailability and ImplementationemeraLD is implemented in C++, and is open source under GPLv3. Source code, documentation, an R interface, and utilities for analysis of summary statistics are freely available at http://github.com/statgen/emeraLD\n\nContactcorbinq@umich.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

UROPA GUI: A web platform for genomic region annotation

The annotation of genomic ranges such as peaks resulting from ChIP-seq/ATAC-seq or other techniques represents a fundamental task of bioinformatics analysis with considerable impact on many downstream analyses. In our previous work, we introduced the Universal Robust Peak Annotator (UROPA), a flexible command line based tool which improves upon the functionality of existing annotation software. In order to reduce the complexity for biologists and clinicians, we have implemented an intuitive web-based graphical user interface (GUI) and fully functional service platform for UROPA. This extension will empower all users to generate annotations for regions of interest interactively.\n\nAvailability and ImplementationThe open source UROPA GUI server was implemented in R Shiny and Python and is available from http://loosolab.mpi-bn.mpg.de. The source code of our App can be downloaded at https://github.molgen.mpg.de/loosolab/UROPA_GUI under the MIT license.

bioinformatics

Gene Set Enrichment Analysis (GSEA) of Upregulated Genes in Cocaine Addiction Reveals miRNAs as Potential Therapeutic Agents

Cocaine addiction is a global health problem that causes substantial damage to the health of addicted individuals around the world. Dopamine synthesizing (DA) neurons in the brain play a vital role in the addiction to cocaine. But the underlying molecular mechanisms that help cocaine exert its addictive effect have not been very well understood. Bioinformatics can be a useful tool in the attempt to broaden our understanding in this area. In the present study, Gene Set Enrichment Analysis (GSEA) was carried out on the upregulated genes from a dataset of DA neurons of post-mortem human brain of cocaine addicts. As a result of this analysis, 3 miRNAs have been identified as having significant influence on transcription of the upregulated genes. These 3 miRNAs hold therapeutic potential for the treatment of cocaine addiction.

bioinformatics

Identification of co-evolving temporal networks

MotivationBiological networks describes the mechanisms which govern cellular functions. Temporal networks show how these networks evolve over time. Studying the temporal progression of network topologies is of utmost importance since it uncovers how a network evolves and how it resists to external stimuli and internal variations. Two temporal networks have co-evolving subnetworks if the topologies of these subnetworks remain similar to each other as the network topology evolves over a period of time. In this paper, we consider the problem of identifying co-evolving pair of temporal networks, which aim to capture the evolution of molecules and their interactions over time. Although this problem shares some characteristics of the well-known network alignment problems, it differs from existing network alignment formulations as it seeks a mapping of the two network topologies that is invariant to temporal evolution of the given networks. This is a computationally challenging problem as it requires capturing not only similar topologies between two networks but also their similar evolution patterns.\n\nResultsWe present an efficient algorithm, Tempo, for solving identifying coevolving subnetworks with two given temporal networks. We formally prove the correctness of our method. We experimentally demonstrate that Tempo scales efficiently with the size of network as well as the number of time points, and generates statistically significant alignments--even when evolution rates of given networks are high. Our results on a human aging dataset demonstrate that Tempo identifies novel genes contributing to the progression of Alzheimers, Huntingtons and Type II diabetes, while existing methods fail to do so.\n\nAvailabilitySoftware is available online (https://www.cise.ufi.edu/[~]relhesha/temporal.zip).\n\nContactrelhesha@ufi.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Benchmarking Statistical Multiple Sequence Alignment

The estimation of multiple sequence alignments of protein sequences is a basic step in many bioinformatics pipelines, including protein structure prediction, protein family identification, and phylogeny estimation. Statistical co-estimation of alignments and trees under stochastic models of sequence evolution has long been considered the most rigorous technique for estimating alignments and trees, but little is known about the accuracy of such methods on biological benchmarks. We report the results of an extensive study evaluating the most popular protein alignment methods as well as the statistical co-estimation method BAli-Phy on 1192 protein data sets from established benchmarks as well as on 120 simulated data sets. Our study (which used more than 230 CPU years for the BAli-Phy analyses alone) shows that BAli-Phy is dramatically more accurate than the other alignment methods on the simulated data sets, but is among the least accurate on the biological benchmarks. There are several potential causes for this discordance, including model misspecification, errors in the reference alignments, and conflicts between structural alignment and evolutionary alignments; future research is needed to understand the most likely explanation for our observations. multiple sequence alignment, BAli-Phy, protein sequences, structural alignment, homology

bioinformatics

Prot-SpaM: Fast alignment-free phylogeny reconstruction based on whole-proteome sequences

Word-based or alignment-free sequence comparison has become an active area of research in bioinformatics. While previous word-frequency approaches calculated rough measures of sequence similarity or dissimilarity, some new alignment-free methods are able to accurately estimate phylogenetic distances between genomic sequences. One of these approaches is Filtered Spaced Word Matches. Herein, we extend this approach to estimate evolutionary distances between complete or incomplete proteomes; our implementation of this approach is called Prot-SpaM. We compare the performance of Prot-SpaM to other alignment-free methods on simulated sequences and on various groups of eukaryotic and prokaryotic taxa. Prot-SpaM can be used to calculate high-quality phylogenetic trees from whole-proteome sequences in a matter of seconds or minutes and often outperforms other alignment-free approaches. The source code of our software is available through Github:\n\nhttps://github.com/jschellh/ProtSpaM

bioinformatics

q2-sample-classifier: machine-learning tools for microbiome classification and regression

Microbiome studies often aim to predict outcomes or differentiate samples based on their microbial compositions, tasks that can be efficiently performed by supervised learning methods. Here we present a benchmark comparison of supervised learning classifiers and regressors implemented in scikit-learn, a Python-based machine-learning library. We additionally present q2-sample-classifier, a plugin for the QIIME 2 microbiome bioinformatics framework, that facilitates application of the scikit-learn classifiers to microbiome data. Random forest, extra trees, and gradient boosting models demonstrate the highest performance for both supervised classification and regression of microbiome data. Automated feature selection and hyperparameter tuning enhance performance of most methods but may not be necessary under all circumstances. The q2-sample-classifier plugin makes these methods more accessible and interpretable to a broad audience of microbiologists, clinicians, and others who wish to utilize supervised learning methods for predicting sample characteristics based on microbiome composition. The q2-sample-classifier source code is available at https://github.com/qiime2/q2-sample-classifier. It is released under a BSD-3-Clause license, and is freely available including for commercial use.

bioinformatics

Quantification of biases in predictions of protein stability changes upon mutations

Bioinformatics tools that predict protein stability changes upon point mutations have made a lot of progress in the last decades and have become accurate and fast enough to make computational mutagenesis experiments feasible, even on a proteome scale. Despite these achievements, they still suffer from important issues that must be solved to allow further improving their performances and utilizing them to deepen our insights into protein folding and stability mechanisms. One of these problems is their bias towards the learning datasets which, being dominated by destabilizing mutations, causes predictions to be better for destabilizing than for stabilizing mutations.\n\nWe thoroughly analyzed the biases in the prediction of folding free energy changes upon point mutations ({Delta}{Delta}G0) and proposed some unbiased solutions. We started by constructing a dataset Ssym of experimentally measured {Delta}{Delta}G0s with an equal number of stabilizing and destabilizing mutations, by collecting mutations for which the structure of both the wild type and mutant protein is available. On this balanced dataset, we assessed the performances of fifteen widely used{Delta}{Delta} G0 predictors. After the astonishing observation that almost all these methods are strongly biased towards destabilizing mutations, especially those that use black-box machine learning, we proposed an elegant way to solve the bias issue by imposing physical symmetries under inverse mutations on the model structure, which we implemented in PoPMuSiCsym. This new predictor constitutes an efficient trade-off between accuracy and absence of biases. Some final considerations and suggestions for further improvement of the predictors are discussed.

bioinformatics

ngsReports: An R Package for managing FastQC reports and other NGS related log files.

MotivationHigh throughput next generation sequencing (NGS) has become exceedingly cheap facilitating studies to be undertaken containing large sample numbers. Quality control (QC) is an essential stage during analytic pipelines and can be found in the outputs of popular bioinformatics tools such as FastQC and Picard. Although these tools provide considerable power when carrying out QC, large sample numbers can make identification of systemic bias a challenge.\n\nResultsWe present ngsReports, an R package designed for the management and visualization of NGS reports from within an R environment. The available methods allow direct import into R of FastQC output as well as that from aligners such as HISAT2, STAR and Bowtie2. Visualization can be carried out across many samples using heatmaps rendered using ggplot2 and plotly. Moreover, these can be displayed in an interactive shiny app or a HTML report. We also provide methods to assess observed GC content in an organism dependent manner for both transcriptomic and genomic datasets. Importantly, hierarchical clustering can be carried out on heatmaps with large sample sizes to quickly identify outliers and batch effects.\n\nAvailability and ImplementationngsReports is available at https://github.com/UofABioinformaticsHub/ngsReports.

bioinformatics

ShinyGO: a graphical enrichment tool for animals and plants

MotivationGene lists are routinely produced from various genome-wide studies. Enrichment analysis can link these gene lists with underlying molecular pathways by using functional categories such as gene ontology (GO).\n\nResultsTo complement existing tools, we developed ShinyGO with several features: (1) large annotation database from GO and many other sources for over 200 plant and animal species, (2) graphical visualization of enrichment results and gene characteristics, and (3) application program interface (API) access to KEGG and STRING for the retrieval of pathway diagrams and protein-protein interaction networks. ShinyGO is an intuitive, graphical web application that can help researchers gain actionable insights from gene lists.\n\nAvailabilityhttp://ge-lab.org/go/\n\nContactgexijin@gmail.com\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Sarek: A portable workflow for whole-genome sequencing analysis of germline and somatic variants

SummaryWhole-genome sequencing (WGS) is a cornerstone of precision medicine, but portable and reproducible open-source workflows for WGS analyses of germline and somatic variants are lacking. We present Sarek, a modular, comprehensive, and easy-to-install workflow, combining a range of software for the identification and annotation of single-nucleotide variants (SNVs), insertion and deletion variants (indels), structural variants, tumor sample heterogeneity, and karyotyping from germline or paired tumor/normal samples. Sarek is implemented in a bioinformatics workflow language (Nextflow) with Docker and Singularity compatible containers, ensuring easy deployment and full reproducibility at any Linux based compute cluster or cloud computing environment. Sarek supports the human reference genomes GRCh37 and GRCh38, and can readily be used both as a core production workflow at sequencing facilities and as a powerful stand-alone tool for individual research groups.\n\nAvailabilitySource code and instructions for local installation are available at GitHub (https://github.com/SciLifeLab/Sarek) under the MIT open-source license, and we invite the research community to contribute additional functionality as a collaborative open-source development project.

bioinformatics

A graph-based algorithm for RNA-seq data normalization

The use of RNA-sequencing has garnered much attention in the recent years for characterizing and understanding various biological systems. However, it remains a major challenge to gain insights from a large number of RNA-seq experiments collectively, due to the normalization problem. Current normalization methods are based on assumptions that fail to hold when RNA-seq profiles become more abundant and heterogeneous. We present a normalization procedure that does not rely on these assumptions, or on prior knowledge about the reference transcripts in those conditions. This algorithm is based on a graph constructed from intrinsic correlations among RNA-seq transcripts and seeks to identify a set of densely connected vertices as references. Application of this algorithm on our benchmark data showed that it can recover the reference transcripts with high precision, thus resulting in high-quality normalization. As demonstrated on a real data set, this algorithm gives good results and is efficient enough to be applicable to real-life data.\n\n2012 ACM Subject ClassificationApplied computing [->] Computational transcriptomics, Applied computing [->] Bioinformatics\n\nDigital Object Identifier10.4230/LIPIcs.WABI.2018.xxx\n\nFundingThis material was based on research supported by the National Heart, Lung, and Blood Institute (NHLBI)-NIH sponsored Programs of Excellence in Glycosciences [grant number HL107152 to B.K.], and partially by NSF [CAREER grant 1350344 to M.M.]. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon.

bioinformatics

GraphSeq: Accelerating String Graph Construction for De Novo Assembly on Spark

Summary: De novo genome assembly is an important application on both uncharacterized genome assembly and variant identification in a reference-unbiased way. In comparison with de Brujin graph, string graph is a lossless data representation for de novo assembly. However, string graph construction is computational intensive. We propose GraphSeq to accelerate string graph construction by leveraging the distributed computing framework.\n\nAvailability and Implementation: GraphSeq is implemented with Scala on Spark and freely available at https://www.atgenomix.com/blog/graphseq.\n\nSupplementary information: Supplementary data are available at Bioinformatics online.

bioinformatics