Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

Privacy-preserving search for chemical compound databases

BackgroundSearching for similar compounds in a database is the most important process for in-silico drug screening. Since a query compound is an important starting point for the new drug, a query holder, who is afraid of the query being monitored by the database server, usually downloads all the records in the database and uses them in a closed network. However, a serious dilemma arises when the database holder also wants to output no information except for the search results, and such a dilemma prevents the use of many important data resources.\n\nResultsIn order to overcome this dilemma, we developed a novel cryptographic protocol that enables database searching while keeping both the query holders privacy and database holders privacy. Generally, the application of cryptographic techniques to practical problems is difficult because versatile techniques are computationally expensive while computationally inexpensive techniques can perform only trivial computation tasks. In this study, our protocol is successfully built only from an additive-homomorphic cryptosystem, which allows only addition performed on encrypted values but is computationally efficient compared with versatile techniques such as general purpose multi-party computation. In an experiment searching ChEMBL, which consists of more than 1,200,000 compounds, the proposed method was 36,900 times faster in CPU time and 12,000 times as efficient in communication size compared with general purpose multi-party computation.\n\nConclusionWe proposed a novel privacy-preserving protocol for searching chemical compound databases. The proposed method, easily scaling for large-scale databases, may help to accelerate drug discovery research by making full use of unused but valuable data that includes sensitive information.

Bioinformatics

PDMQ - Protein Digestion Multi Query software tool to perform in silico digestion of protein/peptide sequences

MotivationIn silico enzymatic digestion tools mostly can be used for digestion of single sequence query, which means a significant limitation in their utility when a number of sequences need to be processed. The other limitation of these applications is the selection options of restriction enzymes that are usually allow only simultaneous digestion. Non-conventional proteins such as cereal prolamins require multi-enzyme multi step digestion, and for cereal proteomics experts this type of application is missing.\n\nResultsPDMQ, Protein Digestion Multi Query application was developed having multi query and multi enzyme options and that way can be customized for any digestion protocol.\n\nAvailability and implementationPDMQ is implemented in C# using the .NET framework and can be downloaded from http://www.agrar.mta.hu/_user/browser/File/bioinformatics/ProteinDigestion_v0_0_0_15.rar

Bioinformatics

NCBI BLAST+ integrated into Galaxy

BackgroundThe NCBI BLAST suite has become ubiquitous in modern molecular biology, used for small tasks like checking capillary sequencing results of single PCR products through to genome annotation or even larger scale pan-genome analyses. For early adopters of the Galaxy web-based biomedical data analysis platform, integrating BLAST was a natural step for sequence comparison workflows.\n\nFindingsThe command line NCBI BLAST+ tool suite was wrapped for use within Galaxy, defining appropriate datatypes as needed, with the goal of making common BLAST tasks easy, and advanced tasks possible.\n\nConclusionsThis effort has been come an informal international collaborative effort, and is deployed and used on Galaxy servers worldwide. Several example use-cases are described herein.

Bioinformatics

Reproducibility Of Parameter Learning With Missing Observations in Naive Wnt Bayesian Network Trained on Normal/Adenomas Samples and Doxycycline Treated LS174T Cell Lines

Insight, Innovation and IntegrationDoxycycline, a derivative of tetracycline, induces gene expression via reversible transcriptional activation. Levels of /3-catenin and other intra/extracellular genetic factors have been influenced in colorectal cancer cell lines, which make doxycycline a potential candidate for cancer chemotherapy. With the aim to build better computational models that show good prediction on test datasets, doxycycline treated cell lines might provide best training samples. This work tests the reproducibility of parameter learning and predictions based on the estimated parameters, using the Naive Bayesian Networks for Wnt pathway in case of missing observations for different nodes. The in silico experiments show the efficacy of causal models as one of the emerging diagnostic tools in development of targeted cancer therapy.\n\nRecent efforts in predicting Wnt signaling activation via inference methods have helped in developing diagnostic models for therapeutic drug targeting. In this manuscript the reproducibility of parameter learning with missing observations in a Bayesian Network and its effect on prediction results for Wnt signaling activation is tested, while training the networks on doxycycline treated LS174T cell lines as well as normal and adenomas samples. This is done in order to check the effectiveness of using Bayesian Network as a tool for modeling Wnt pathway when certain observations are missing. Experimental analysis suggest that prediction results are reproducible with negligible deviations. Anomalies in estimated parameters are accounted for due to the Bayesian Network model. Also, an interesting case regarding usage of hypothesis testing came up while proving the statistical significance of different design setups of the BN model which was trained on the same data. It was found that hypothesis testing may not be the correct way to check the significance between design setups for the aforementioned case, especially when the structure of the model is same. Finally, in comparison to the biologically inspired models, the naive bayesian model may give accurate results but this accuracy comes at the cost of loss of crucial biological knowledge which might help reveal hidden relations among intra/extracellular factors affecting the Wnt pathway.

Bioinformatics

Alternative splicing QTLs in European and African populations using Altrans, a novel method for splice junction quantification

With the advent of RNA-sequencing technology we now have the power to detect different types of alternative splicing and how DNA variation affects splicing. However, given the short read lengths used in most population based RNA-sequencing experiments, quantifying transcripts accurately remains a challenge. Here we present a novel method, Altrans, for discovery of alternative splicing quantitative trait loci (asQTLs). To assess the performance of Altrans we compared it to Cufflinks, a well-established transcript quantification method. Simulations show that in the presence of transcripts absent from the annotation, Altrans performs better in quantifications than Cufflinks. We have applied Altrans and Cufflinks to the Geuvadis dataset, which comprises samples from European and African populations, and discovered (FDR = 1%) 1806 and 243 asQTLs with Altrans, and 1596 and 288 asQTLs with Cufflinks for Europeans and Africans, respectively. Although Cufflinks results replicated better across the two populations, this likely due to the increased sensitivity of Altrans in detecting harder to detect associations. We show that, by discovering a set of asQTLs in a smaller subset of European samples and replicating these in the remaining larger subset of Europeans, both methods achieve similar replication levels (94% and 98% replication in Altrans and Cufflinks, respectively). We find that method specific asQTLs are largely due to different types of alternative splicing events detected by each method. We overlapped the asQTLs with biochemically active regions of the genome and observed significant enrichments for many functional marks and variants in splicing regions, highlighting the biological relevance of the asQTLs identified. All together, we present a novel approach for discovering asQTLs that is a more direct assessment of splicing compared to other methods and is complementary to other transcript quantification methods.

Bioinformatics

Segmentation of Noisy Signals Generated By a Nanopore

Nanopore-based single-molecule sequencing techniques exploit ionic current steps produced as biomolecules pass through a pore to reconstruct properties of the sequence. A key task in analyzing complex nanopore data is discovering the boundaries between these steps, which has traditionally been done in research labs by hand. We present an automated method of analyzing nanopore data, by detecting regions of ionic current corresponding to the translocation of a biomolecule, and then segmenting the region. The segmenter uses a divide-and-conquer method to recursively discover boundary points, with an implementation that works several times faster than real time and that can handle low-pass filtered signals.

Bioinformatics

Linked annotations: a middle ground for manual curation of biomedical databases and text corpora

Annotators of text corpora and biomedical databases carry out the same labor-intensive task to manually extract structured data from unstructured text. Tasks are needlessly repeated because text corpora are widely scattered. We envision that a linked annotation resource unifying many corpora could be a game changer. Such an open forum will help focus on novel annotations and on optimally benefiting from the energy of many experts. As proof-of-concept, we annotated protein subcellular localization in 100 abstracts cited by UniProtKB. The detailed comparison between our new corpus and the original UniProtKB annotations revealed sustained novel annotations for 42% of the entries (proteins). In a unified linked annotation resource these could immediately extend the utility of text corpora beyond the text-mining community. Our example motivates the central idea that linked annotations from text corpora can complement database annotations.

Bioinformatics

Do Read Errors Matter for Genome Assembly?

While most current high-throughput DNA sequencing technologies generate short reads with low error rates, emerging sequencing technologies generate long reads with high error rates. A basic question of interest is the tradeoff between read length and error rate in terms of the information needed for the perfect assembly of the genome. Using an adversarial erasure error model, we make progress on this problem by establishing a critical read length, as a function of the genome and the error rate, above which perfect assembly is guaranteed. For several real genomes, including those from the GAGE dataset, we verify that this critical read length is not significantly greater than the read length required for perfect assembly from reads without errors.

Bioinformatics

Sincell: Bioconductor package for the statistical assessment of cell-state hierarchies from single-cell RNA-seq data

SummaryCell differentiation processes are achieved through a continuum of hierarchical intermediate cell-states that might be captured by single-cell RNA seq. Existing computational approaches for the assessment of cell-state hierarchies from single-cell data might be formalized under a general framework composed of i) a metric to assess cell-to-cell similarities (combined or not with a dimensionality reduction step), and ii) a graph-building algorithm (optionally making use of a cells-clustering step). Sincell R package implements a methodological toolbox allowing flexible workflows under such framework. Furthermore, Sincell contributes new algorithms to provide cell-state hierarchies with statistical support while accounting for stochastic factors in single-cell RNA seq. Graphical representations and functional association tests are provided to interpret hierarchies. Sincell functionalities are illustrated in a real case study where its ability to discriminate noisy from stable cell-state hierarchies is demonstrated.\n\nAvailability and implementationSincell is an open-source R/Bioconductor package available at http://bioconductor.org/packages/3.1/bioc/html/sincell.html. A detailed manual and vignette describing functions and workflows is provided with the package.\n\nContact: antonio.rausell@isb-sib.ch

Bioinformatics

Assembly by Reduced Complexity (ARC): a hybrid approach for targeted assembly of homologous sequences.

Analysis of High-throughput sequencing (HTS) data is a difficult problem, especially in the context of non-model organisms where comparison of homologous sequences may be hindered by the lack of a close reference genome. Current mapping-based methods rely on the availability of a highly similar reference sequence, whereas de novo assemblies produce anonymous (unannotated) contigs that are not easily compared across samples. Here, we present Assembly by Reduced Complexity (ARC) a hybrid mapping and assembly approach for targeted assembly of homologous sequences. ARC is an open-source project (http://ibest.github.io/ARC/) implemented in the Python language and consists of the following stages: 1) align sequence reads to reference targets, 2) use alignment results to distribute reads into target specific bins, 3) perform assemblies for each bin (target) to produce contigs, and 4) replace previous reference targets with assembled contigs and iterate. We show that ARC is able to assemble high quality, unbiased mitochondrial genomes seeded from 11 progressively divergent references, and is able to assemble full mitochondrial genomes starting from short, poor quality ancient DNA reads. We also show ARC compares favorably to de novo assembly of a large exome capture dataset for CPU and memory requirements; assembling 7,627 individual targets across 55 samples, completing over 1.3 million assemblies in less than 78 hours, while using under 32 Gb of system memory. ARC breaks the assembly problem down into many smaller problems, solving the anonymous contig and poor scaling inherent in some de novo assembly methods and reference bias inherent in traditional read mapping.

Bioinformatics

dbSUPER: a database of super-enhancers in mouse and human genome

Super-enhancers are the clusters of transcriptional enhancers that can drive cell-type-specific gene expression and also crucial in cell identity. Many disease-associated sequence variations are enriched in super-enhancer regions of disease-relevant cell types. Thus, super-enhancers can be used as potential biomarkers for disease diagnosis and therapeutics. Current studies have identified super-enhancers in more than 100 cell types and demonstrated their functional importance. However, no centralized resource to integrate all these findings is available yet. We developed dbSUPER (http://bioinfo.au.tsinghua.edu.cn/dbsuper/), the first integrated and interactive database of super-enhancers, with the primary goal of providing a resource for assistance in further studies related to transcriptional control of cell identity and disease. dbSUPER provides a responsive and user-friendly web interface to facilitate efficient and comprehensive search and browsing. The data can be easily sent to Galaxy instances, GREAT and Cistrome web servers for downstream analysis, and can be visualized in UCSC genome browser while custom tracks added automatically. The data can be downloaded and exported in variety of formats. Further, dbSUPER lists genes associated with the super-enhancers and links to various other databases such as GeneCards, UniProt and Entrez. dbSUPER also provides an overlap analysis tool, to annotate user defined regions. We believe dbSUPER is a valuable resource for the biologists and genetic research communities.

Bioinformatics

Histoimmunogenetics Markup Language 1.0: Reporting Next Generation Sequencing-based HLA and KIR Genotyping

We present an electronic format for exchanging data for HLA and KIR genotyping with extensions for next-generation sequencing (NGS). This format addresses NGS data exchange by refining the Histoimmunogenetics Markup Language (HML) to conform to the proposed Minimum Information for Reporting Immunogenomic NGS Genotyping (MIRING) reporting guidelines (miring.immunogenomics.org). Our refinements of HML include two major additions. First, NGS is supported by new XML structures to capture additional NGS data and metadata required to produce a genotyping result, including analysis-dependent (dynamic) and method-dependent (static) components. A full genotype, consensus sequence, and the surrounding metadata are included directly, while the raw sequence reads and platform documentation are externally referenced. Second, genotype ambiguity is fully represented by integrating Genotype List Strings, which use a hierarchical set of delimiters to represent allele and genotype ambiguity in a complete and accurate fashion. HML also continues to enable the transmission of legacy methods (e.g. site-specific oligonucleotide, sequence-specific priming, and sequence based typing (SBT)), adding features such as allowing multiple group-specific sequencing primers, and fully leveraging techniques that combine multiple methods to obtain a single result, such as SBT integrated with NGS.\n\nAbbreviations

Bioinformatics

Relationship between tumor grade and geometrical complexity in prostate cancer

Prostate cancer exhibits high mathematical complexity due to the disruption of tissue architecture. An important part of the diagnostic of prostate tumor samples is the histological evaluation of cellular and glandular organization. The Gleason grade and score, a commonly used prognostic indicator of patient outcome, is based on the match of glandular architectural patterns with standard patterns. Unfortunately, the subjective nature of visual grading leads to variations in scoring by different pathologists. We proposed the fractal dimension of the lumen and the Lempel-Zip complexity of the histopathological patterns as useful descriptors aiding pathologist to standardize histological classification and thus prognosis and therapy planning.\n\nHighlightsO_LIgeometrical complexity of prostate cancer\nC_LI

Bioinformatics

Using Mixtures of Biological Samples as Process Controls for RNA-sequencing experiments

BackgroundGenome-scale \"-omics\" measurements are challenging to benchmark due to the enormous variety of unique biological molecules involved. Mixtures of previously-characterized samples can be used to benchmark repeatability and reproducibility using component proportions as truth for the measurement. We describe and evaluate experiments characterizing the performance of RNA-sequencing (RNA-Seq) measurements, and discuss cases where mixtures can serve as effective process controls.\n\nResultsWe apply a linear model to total RNA mixture samples in RNA-seq experiments. This model provides a context for performance benchmarking. The parameters of the model fit to experimental results can be evaluated to assess bias and variability of the measurement of a mixture. A linear model describes the behavior of mixture expression measures and provides a context for performance benchmarking. Residuals from fitting the model to experimental data can be used as a metric for evaluating the effect that an individual step in an experimental process has on the linear response function and precision of the underlying measurement while identifying signals affected by interference from other sources. Effective benchmarking requires well-defined mixtures, which for RNA-Seq requires knowledge of the messenger RNA (mRNA) content of the individual total RNA components. We demonstrate and evaluate an experimental method suitable for use in genome-scale process control and lay out a method utilizing spike-in controls to determine mRNA content of total RNA in samples.\n\nConclusionsGenome-scale process controls can be derived from mixtures. These controls relate prior knowledge of individual components to a complex mixture, allowing assessment of measurement performance. The mRNA fraction accounts for differential enrichment of mRNA from varying total RNA samples. Spike-in controls can be utilized to measure this relationship between mRNA content and input total RNA. Our mixture analysis method also enables estimation of the proportions of an unknown mixture, even when component-specific markers are not previously known, whenever pure components are measured alongside the mixture.

Bioinformatics

Learning Immune-Defectives Graph through Group Tests

This paper deals with an abstraction of a unified problem of drug discovery and pathogen identification. Here, the \"lead compounds\" are abstracted as inhibitors, pathogenic proteins as defectives, and the mixture of \"ineffective\" chemical compounds and non-pathogenic proteins as normal items. A defective could be immune to the presence of an inhibitor in a test. So, a test containing a defective is positive iff it does not contain its \"associated\" inhibitor. The goal of this paper is to identify the defectives, inhibitors, and their \"associations\" with high probability, or in other words, learn the Immune Defectives Graph (IDG). We propose a probabilistic non-adaptive pooling design, a probabilistic two-stage adaptive pooling design and decoding algorithms for learning the IDG. For the two-stage adaptive-pooling design, we show that the sample complexity of the number of tests required to guarantee recovery of the inhibitors, defectives and their associations with high probability, i.e., the upper bound, exceeds the proposed lower bound by a logarithmic multiplicative factor in the number of items. For the non-adaptive pooling design, in the large inhibitor regime, we show that the upper bound exceeds the proposed lower bound by a logarithmic multiplicative factor in the number of inhibitors.

Bioinformatics

Discovery of large genomic inversions using pooled clone sequencing

MotivationThere are many different forms of genomic structural variation that can be broadly classified into two groups as copy number variation (CNV) and balanced rearrangements. Although many algorithms are now available in the literature that aim to characterize CNVs, discovery of balanced rearrangements (inversions and translocations) remains an open problem. This is mainly because the breakpoints of such events typically lie within segmental duplications and common repeats, which reduce the mappability of short reads. The 1000 Genomes Project spearheaded the development of several methods to identify inversions, however, they are limited to relatively short inversions, and there are currently no available algorithms to discover large inversions using high throughput sequencing technologies (HTS).\n\nResultsHere we propose to use a sequencing method (Kitzman et al., 2011) originally developed to improve haplotype phasing to characterize large genomic inversions. This method, called pooled clone sequencing, merges the advantages of clone based sequencing approaches with the speed and cost efficiency of HTS technologies. Using data generated with pooled clone sequencing method, we developed a novel algorithm, dipSeq, to discover large inversions (>500 Kbp). We show the power of dipSeq first on simulated data, and then apply it to the genome of a HapMap individual (NA12878). We were able to accurately discover all previously known and experimentally validated large inversions in the same genome. We also identified a novel inversion, and confirmed using fluorescent in situ hybridization.\n\nAvailabilityImplementation of the dipSeq algorithm is available at https://github.com/BilkentCompGen/dipseq\n\nContactcalkan@cs.bilkent.edu.tr, francesca.antonacci@uniba.it

Bioinformatics

A multi-method approach for proteomic network inference in 11 human cancers

Protein expression and post-translational modification levels are tightly regulated in neoplastic cells to maintain cellular processes known as cancer hallmarks. The first Pan-Cancer initiative of The Cancer Genome Atlas (TCGA) Research Network has aggregated protein expression profiles for 3,467 patient samples from 11 tumor types using the antibody based reverse phase protein array (RPPA) technology. The resultant proteomic data can be utilized to computationally infer protein-protein interaction (PPI) networks and to study the commonalities and differences across tumor types. In this study, we compare the performance of 13 established network inference methods in their capacity to retrieve literature-curated pathway interactions from RPPA data. We observe that no single method has the best performance in all tumor types, but a group of six methods, including diverse techniques such as correlation, mutual information, and regression, consistently rank highly among the tested methods. A consensus network from this high-performing group reveals that signal transduction events involving receptor tyrosine kinases (RTKs), the RAS/MAPK pathway, and the PI3K/AKT/mTOR pathway, as well as innate and adaptive immunity signaling, are the most significant PPIs shared across all tumor types. Our results illustrate the utility of the RPPA platform as a tool to study proteomic networks in cancer.\n\nAvailabilityPPI networks from the TCGA or user-provided data can be visualized with the ProtNet web application at http://www.sanderlab.org/protnet/.

Bioinformatics

A GENE FEATURE ENUMERATION APPROACH FOR DESCRIBING HLA ALLELE POLYMORPHISM

HLA genotyping via next generation sequencing (NGS) poses challenges for the use of HLA allele names to analyze and discuss sequence polymorphism. NGS will identify many new synonymous and non-coding HLA sequence variants. Allele names identify the types of nucleotide polymorphism that define an allele (non-synonymous, synonymous and non-coding changes), but do not describe how polymorphism is distributed among the individual features (the flanking untranslated regions, exons and introns) of a gene. Further, HLA alleles cannot be named in the absence of antigen-recognition domain (ARD) encoding exons. Here, a system for describing HLA polymorphism in terms of HLA gene features (GFs) is proposed. This system enumerates the unique nucleotide sequences for each GF in an HLA gene, and records these in a GF enumeration notation that allows both more granular dissection of allele-level HLA polymorphism, and the discussion and analysis of GFs in the absence of ARD-encoding exon sequences.\n\nAbbreviations

Bioinformatics