Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Bio-Docklets: Virtualization Containers for Single-Step Execution of NGS Pipelines.

BackgroundProcessing of Next-Generation Sequencing (NGS) data requires significant technical skills, involving installation, configuration, and execution of bioinformatics data pipelines, in addition to specialized post-analysis visualization and data mining software. In order to address some of these challenges, developers have leveraged virtualization containers, towards seamless deployment of preconfigured bioinformatics software and pipelines on any computational platform.\n\nFindingsWe present an approach for abstracting the complex data operations of multi-step, bioinformatics pipelines for NGS data analysis. As examples, we have deployed two pipelines for RNAseq and CHIPseq, pre-configured within Docker virtualization containers we call Bio-Docklets. Each Bio-Docklet exposes a single data input and output endpoint and from a user perspective, running the pipelines is as simple as running a single bioinformatics tool. This is achieved through a \"meta-script\" that automatically starts the Bio-Docklets, and controls the pipeline execution through the BioBlend software library and the Galaxy Application Programming Interface (API). The pipelne output is post-processed using the Visual Omics Explorer (VOE) framework, providing interactive data visualizations that users can access through a web browser.\n\nConclusionsThe goal of our approach is to enable easy access to NGS data analysis pipelines for nonbioinformatics experts, on any computing environment whether a laboratory workstation, university computer cluster, or a cloud service provider. Besides end-users, the Bio-Docklets also enables developers to programmatically deploy and run a large number of pipeline instances for concurrent analysis of multiple datasets.

bioinformatics

Reproducible Bioconductor Workflows Using Browser-Based Interactive Notebooks And Containers

ObjectiveBioinformatics publications typically include complex software workflows that are difficult to describe in a manuscript. We describe and demonstrate the use of interactive software notebooks to document and distribute bioinformatics research. We provide a user-friendly tool, BiocImageBuilder, to allow users to easily distribute their bioinformatics protocols through interactive notebooks uploaded to either a GitHub repository or a private server.\n\nMaterials and methodsWe present three different interactive Jupyter notebooks using R and Bioconductor workflows to infer differential gene expression, analyze cross-platform datasets and process RNA-seq data. These interactive notebooks are available on GitHub. The analytical results can be viewed in a browser. Most importantly, the software contents can be executed and modified. This is accomplished using Binder, which runs the notebook inside software containers, thus avoiding the need for installation of any software and ensuring reproducibility. All the notebooks were produced using custom files generated by BiocImageBuilder.\n\nResultsBiocImageBuilder facilitates the publication of workflows with a point-and-click user interface. We demonstrate that interactive notebooks can be used to disseminate a wide range of bioinformatics analyses. The use of software containers to mirror the original software environment ensures reproducibility of results. Parameters and code can be dynamically modified, allowing for robust verification of published results and encouraging rapid adoption of new methods.\n\nConclusionGiven the increasing complexity of bioinformatics workflows, we anticipate that these interactive software notebooks will become as ubiquitous and necessary for documenting software methods as traditional laboratory notebooks have been for documenting bench protocols.

bioinformatics

Features of ChIP-seq data peak calling algorithms with good operating characteristics

Author descriptionReuben Thomas is a Staff Research Scientist in the Bioinformatics Core at Gladstone Institutes\n\nSean Thomas is a Staff Research Scientist in the Bioinformatics Core at Gladstone Institutes\n\nAlisha K Holloway is the Director of Bioinformatics at Phylos Biosciences, visiting scientist at Gladstone Institutes and Adjunct Assistant Professor in Biostatistics at the University of California, San Francisco.\n\nKatherine S Pollard is a Senior Investigator at Gladstone Institutes and Professor of Biostatistics at University of California, San Francisco.\n\nKey PointsO_LIPeak-calling using Chip-seq data consists of two sub-problems: identifying candidate peaks and testing candidate peaks for statistical significance.\nC_LIO_LITwelve features of the two sub-problems of peak-calling methods are identified.\nC_LIO_LIMethods that explicitly combine the signals from ChIP and input samples are less powerful than methods that do not.\nC_LIO_LIMethods that use windows of different sizes to scan the genome for potential peaks are more powerful than ones that do not.\nC_LIO_LIMethods that use a Poisson test to rank their candidate peaks are more powerful than those that use a Binomial test.\nC_LI\n\nAbstractChromatin immunoprecipitation followed by sequencing (ChIP-seq) is an important tool for studying gene regulatory proteins, such as transcription factors and histones. Peak calling is one of the first steps in analysis of these data. Peak-calling consists of two sub-problems: identifying candidate peaks and testing candidate peaks for statistical significance. We surveyed 30 methods and identified 12 features of the two sub-problems that distinguish methods from each other. We picked six methods (GEM, MACS2, MUSIC, BCP, TM and ZINBA) that span this feature space and used a combination of 300 simulated ChIP-seq data sets, 3 real data sets and mathematical analyses to identify features of methods that allow some to perform better than others. We prove that methods that explicitly combine the signals from ChIP and input samples are less powerful than methods that do not. Methods that use windows of different sizes are more powerful than ones that do not. For statistical testing of candidate peaks, methods that use a Poisson test to rank their candidate peaks are more powerful than those that use a Binomial test. BCP and MACS2 have the best operating characteristics on simulated transcription factor binding data. GEM has the highest fraction of the top 500 peaks containing the binding motif of the immunoprecipitated factor, with 50% of its peaks within 10 base pairs (bp) of a motif. BCP and MUSIC perform best on histone data. These findings provide guidance and rationale for selecting the best peak caller for a given application.

Bioinformatics

Duplicates, redundancies, and inconsistencies in the primary nucleotide databases: a descriptive study

GenBank, the EMBL European Nucleotide Archive, and the DNA DataBank of Japan, known collectively as the International Nucleotide Sequence Database Collaboration or INSDC, are the three most significant nucleotide sequence databases. Their records are derived from laboratory work undertaken by different individuals, by different teams, with a range of technologies and assumptions, and over a period of decades. As a consequence, they contain a great many duplicates, redundancies, and inconsistencies, but neither the prevalence nor the characteristics of various types of duplicates have been rigorously assessed. Existing duplicate detection methods in bioinformatics only address specific duplicate types, with inconsistent assumptions; and the impact of duplicates in bioinformatics databases has not been carefully assessed, making it difficult to judge the value of such methods. Our goal is to assess the scale, kinds, and impact of duplicates in bioinformatics databases, through a retrospective analysis of merged groups in INSDC databases. Our outcomes are threefold: (1) We analyse a benchmark dataset consisting of duplicates manually identified in INSDC - a dataset of 67,888 merged groups with 111,823 duplicate pairs across 21 organisms from INSDC databases - in terms of the prevalence, types, and impacts of duplicates. (2) We categorise duplicates at both sequence and annotation level, with supporting quantitative statistics, showing that different organisms have different prevalence of distinct kinds of duplicate. (3) We show that the presence of duplicates has practical impact via a simple case study on duplicates, in terms of GC content and melting temperature. We demonstrate that duplicates not only introduce redundancy, but can lead to inconsistent results for certain tasks. Our findings lead to a better understanding of the problem of duplication in biological databases.

bioinformatics

SNVPhyl: A Single Nucleotide Variant Phylogenomics pipeline for microbial genomic epidemiology

MotivationThe recent widespread application of whole-genome sequencing (WGS) for microbial disease investigations has spurred the development of new bioinformatics tools, including a notable proliferation of phylogenomics pipelines designed for infectious disease surveillance and outbreak investigation. Transitioning the use of WGS data out of the research lab and into the front lines of surveillance and outbreak response requires user-friendly, reproducible, and scalable pipelines that have been well validated.\n\nResultsSNVPhyl (Single Nucleotide Variant Phylogenomics) is a bioinformatics pipeline for identifying high-quality SNVs and constructing a whole genome phylogeny from a collection of WGS reads and a reference genome. Individual pipeline components are integrated into the Galaxy bioinformatics framework, enabling data analysis in a user-friendly, reproducible, and scalable environment. We show that SNVPhyl can detect SNVs with high sensitivity and specificity and identify and remove regions of high SNV density (indicative of recombination). SNVPhyl is able to correctly distinguish outbreak from non-outbreak isolates across a range of variant-calling settings, sequencing-coverage thresholds, or in the presence of contamination.\n\nAvailabilitySNVPhyl is available as a Galaxy workflow, Docker and virtual machine images, and a Unix-based command-line application. SNVPhyl is released under the Apache 2.0 license and available at http://snvphyl.readthedocs.io/ or at https://github.com/phac-nml/snvphyl-galaxy.

bioinformatics

DepEst: an R package of important dependency estimators for gene network inference algorithms

Gene network inference algorithms (GNI) are popular in bioinformatics area. In almost all GNI algorithms, the main process is to estimate the dependency (association) scores among the genes of the dataset.\n\nWe present a bioinformatics tool, DepEst (Dependency Estimators), which is a powerful and flexible R package that includes 11 important dependency score estimators that can be used in almost all GNI Algorithms. DepEst is the first bioinformatics package that includes such a large number of estimators that runs both in parallel and serial.\n\nDepEst is currently available at https://github.com/altayg/Depest. Package access link, instructions, various workflows and example data sets are provided in the supplementary file.

bioinformatics

TeachEnG: a Teaching Engine for Genomics

MotivationBioinformatics is a rapidly growing field that has emerged from the synergy of computer science, statistics, and biology. Given the interdisciplinary nature of bioinformatics, many students from diverse fields struggle with grasping bioinformatic concepts only from classroom lectures. Interactive tools for helping students reinforce their learning would be thus desirable. Here, we present an interactive online educational tool called TeachEnG (acronym for Teaching Engine for Genomics) for reinforcing key concepts in sequence alignment and phylogenetic tree reconstruction. Our instructional games allow students to align sequences by hand, fill out the dynamic programming matrix in the Needleman-Wunsch global sequence alignment algorithm, and reconstruct phylogenetic trees via the maximum parsimony and Unweighted Pair Group Method with Arithmetic mean (UPGMA) algorithms. With an easily accessible interface and instant visual feedback, TeachEnG will help promote active learning in bioinformatics.\n\nAvailability and Implementation: TeachEnG is freely available at http://song.igb.illinois.edu/TeachEnG/. It is written in JavaScript and compatible with Firefox, Safari, Chrome, and Microsoft Edge.\n\nContact: songi@illinois.edu

bioinformatics

Accurate And Fast Feature Selection Workflow For High-Dimensional Omics Data

We are moving into the age of Big Data in biomedical research and bioinformatics. This trend could be encapsulated in this simple formula: D = S x F, where the volume of data generated (D) increases in both dimensions: the number of samples (S) and the number of sample features (F). Frequently, a typical bioinformatics problem (e.g. classification) includes redundant and irrelevant features that can result, in the worst-case scenario, in false positive results. Then, Feature Selection (FS) constitutes an enormous challenge. Despite the number and diversity of algorithms available, the proper choice of an approach for facing a specific problem often falls in a grey zone. In this study, we select a subset of FS methods to develop an efficient workflow and an R package for bioinformatics machine learning problems. We cover relevant issues concerning FS, ranging from domains problems to algorithm solutions and computational tools. Finally, we use seven different proteomics and gene expression datasets to evaluate the workflow and guide the FS process.

bioinformatics

A rapid and accurate approach for prediction of interactomes from co-elution data (PrInCE)

BackgroundAn organisms protein interactome, or complete network of protein-protein interactions, defines the protein complexes that drive cellular processes. Techniques for studying protein complexes have traditionally applied targeted strategies such as yeast two-hybrid or affinity purification-mass spectrometry to assess protein interactions. However, given the vast number of protein complexes, more scalable methods are necessary to accelerate interaction discovery and to construct whole interactomes. We recently developed a complementary technique based on the use of protein correlation profiling (PCP) and stable isotope labeling in amino acids in cell culture (SILAC) to assess chromatographic co-elution as evidence of interacting proteins. Importantly, PCP-SILAC is also capable of measuring protein interactions simultaneously under multiple biological conditions, allowing the detection of treatment-specific changes to an interactome. Given the uniqueness and high dimensionality of co-elution data, new tools are needed to compare protein elution profiles, control false discovery rates, and construct an accurate interactome.\n\nResultsHere we describe a freely available bioinformatics pipeline, PrInCE, for the analysis of co-elution data. PrInCE is a modular, open-source library that is computationally inexpensive, able to use label and label-free data, and capable of detecting tens of thousands of protein-protein interactions. Using a machine learning approach, PrInCE offers greatly reduced run time, better performance, prediction of protein complexes, and greater ease of use over previous bioinformatics tools for co-elution data. PrInCE is implemented in Matlab (version R2015b). Source code and standalone executable programs for Windows and Mac OSX are available at https://github.com/fosterlab/PrInCE, where usage instructions can be found. An example dataset and output are also provided for testing purposes.\n\nConclusionsPrInCE is the first fast and easy-to-use data analysis pipeline that predicts interactomes and protein complexes from co-elution data. PrInCE allows researchers without bioinformatics proficiency to analyze high-throughput co-elution datasets.

bioinformatics

unitas: the universal tool for annotation of small RNAs

BackgroundNext generation sequencing is a key technique in small RNA biology research that has led to the discovery of functionally different classes of small non-coding RNAs in the past years. However, reliable annotation of the extensive amounts of small non-coding RNA data produced by high-throughput sequencing is time-consuming and requires robust bioinformatics expertise. Moreover, existing tools have a number of shortcomings including a lack of sensitivity under certain conditions, limited number of supported species or detectable sub-classes of small RNAs.\n\nResultsHere we introduce unitas, an out-of-the-box ready software for complete annotation of small RNA sequence datasets, supporting the wide range of species for which non-coding RNA reference sequences are available in the Ensembl databases (currently more than 800). unitas combines high quality annotation and numerous analysis features in a user-friendly manner. A complete annotation can be started with one simple shell command, making unitas particularly useful for researchers not having access to a bioinformatics facility. Noteworthy, the algorithms implemented in unitas are on par or even outperform comparable existing tools for small RNA annotation that base on available ncRNA sequence information.\n\nConclusionsunitas brings together annotation and analysis features that hitherto required the installation of numerous different bioinformatics tools which can pose a challenge for the non-expert user. With this, unitas overcomes the problem of read normalization. Moreover, the high quality of sequence annotation and analysis, paired with the ease of use, make unitas a valuable tool for researchers in all fields connected to small RNA biology.

bioinformatics

canvasXpress: A versatile interactive high-resolution scientific multi-panel visualization toolkit

To the Editor: CanvasXpress (https://canvasxpress.org) was developed as the core visualization component for bioinformatics and systems biology analysis at Bristol-Myers Squibb and further enhanced by scientists around the world and served as a key visualization engine for many popular bioinformatics tools1,2,3,4,5,6. It offers a rich set of interactive plots to display scientific and genomics data, such as oncoprint of cancer mutations, heatmap, 3D scatter, violin, radar, and profile plots (Figure 1, canvasXpress plots arranged by canvasDesigner https://baohongz.github.io/canvasDesigner). Recently, the reproducibility and usability of the package in real world bioinformatics and clinical use cases have been improved significantly witnessed by continuous add-on features and wide adoption of the toolkit in the scientific communities. Furthermore, It is the first noteworthy package harmonizing real time interactive exploring and analyzing of big data, full-fledged customizing of look-n-feel, and producing multi-panel publication-ready figures in PDF format simultaneously.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=84 SRC=\"FIGDIR/small/186213_fig1.gif\" ALT=\"Figure 1\">\nView larger version (36K):\norg.highwire.dtl.DTLVardef@1d8392borg.highwire.dtl.DTLVardef@916a49org.highwire.dtl.DTLVardef@d90b05org.highwire.dtl.DTLVardef@162a8fc_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Versatile plots generated by canvasXpress. A) Oncoprint of cancer mutations; B) Gene expression heatmap with comprehensive annotations; C) 3D scatter plot; D) Violin plot; E) Radar plot with annotations; F) Profile plot of gene expression.\n\nC_FIG

bioinformatics

Clustering of Circular Consensus Sequences: Accurate Error Correction and Assembly of Single Molecule Real-Time Reads from Multiplexed Amplicon Libraries

BACKGROUNDTargeted resequencing with high-throughput sequencing (HTS) platforms can be used to efficiently interrogate the genomes of large numbers of individuals. A critical challenge for research and applications using HTS data, especially from long-read platforms, is errors arising from technological limits and bioinformatic algorithms.\n\nRESULTSA single molecule real-time (SMRT) sequencing-error correction and assembly pipeline, C3S-LAA, was developed for libraries of pooled amplicons. By uniquely leveraging the structure of SMRT sequence data (comprised of multiple low quality subreads from which higher quality circular consensus sequences are formed) to cluster raw reads, C3S-LAA produced accurate consensus sequences and assemblies of overlapping amplicons from single sample and multiplexed libraries. In contrast, despite read depths in excess of 100X per amplicon, the standard long amplicon analysis module from Pacific Biosciences generated unexpected numbers of amplicon sequences with substantial inaccuracies in the consensus sequences. A bootstrap analysis showed that the C3S-LAA pipeline per se was effective at removing bioinformatic sources of error, but in rare cases a read depth of nearly 400X was not sufficient to overcome minor but systematic errors inherent to amplification or sequencing.\n\nCONCLUSIONSC3S-LAA uses a novel processing algorithm for SMRT amplicon-sequence data that produces accurate consensus sequences and local sequence assemblies. The community standard long amplicon analysis module from Pacific Biosciences is prone to substantial errors that raise concerns about findings based on this pipeline. The method developed here removed this confounding bioinformatics source of error, allowing for the identification of limited instances of errors due to DNA amplification or sequencing.

bioinformatics

De novo profile generation based on sequence context specificity with the long short-term memory network

Long short-term memory (LSTM) is one of the most attractive deep learning methods to learn time series or contexts of input data. Increasing studies, including biological sequence analyses in bioinformatics, utilize this architecture. Amino acid sequence profiles are widely used for bioinformatics studies, such as sequence similarity searches, multiple alignments, and evolutionary analyses. Currently, many biological sequences are becoming available, and the rapidly increasing amount of sequence data emphasizes the importance of scalable generators of amino acid sequence profiles. We employed the LSTM network and developed a novel profile generator to construct profiles without any assumptions, except for input sequence context. Our method could generate better profiles than existing de novo profile generators, including CSBuild and RPS-BLAST, on the basis of profile-sequence similarity search performance with linear calculation costs against input sequence size. In addition, we analyzed the effects of the memory power of LSTM and found that LSTM had high potential power to detect long-range interactions between amino acids, as in the case of beta-strand formation, which has been a difficult problem in protein bioinformatics using sequence information. We demonstrated the importance of sequence context and the feasibility of LSTM on biological sequence analyses. Our results demonstrated the effectiveness of memories in LSTM and showed that our de novo profile generator, SPBuild, achieved higher performance than that of existing methods for profile prediction of beta-strands, where long-range interactions of amino acids are important and are known to be difficult for the existing window-based prediction methods. Our findings will be useful for the development of other prediction methods related to biological sequences by machine learning methods.

bioinformatics

PEPOP: new approaches to mimic non-continous epitopes

Bioinformatics methods are helpful to identify new molecules for diagnostic or therapeutic applications. For example, the use of peptides capable of mimicking binding sites has several benefits as replacing a protein difficult to produce, or toxic. Using peptides is less expensive. Peptides are easier to manipulate, and can be used as drugs. Continuous epitope predicted by bioinformatics tools are commonly used and these sequential epitopes are used as such in further experiments. Numerous discontinuous epitope predictors have been developed but only two bioinformatics tools proposed so far to predict peptide sequences: Superficial and PEPOP. PEPOP can generate series of peptide sequences that can replace continuous or discontinuous epitopes in their interaction with their cognate antibody. We have developed an improved version of PEPOP dedicated to answer to the experimentalists need for a tool able to handle proteins and to turn them into peptides. The PEPOP web site has been reorganized by peptide prediction category and is therefore better formulated to experimental designs. Since the first version of PEPOP, 32 new methods of peptide design were developed. In total, PEPOP proposes 35 methods in which 34 deal specifically with discontinuous epitopes, the most represented epitope type in nature.\n\nWe present the user-friendly, well-structured web-site of PEPOP and its validation through the use of predicted immunogenic or antigenic peptides mimicking discontinuous epitopes in different experimental ways. PEPOP proposes 35 methods of peptide design to guide experimentalists in using peptides potentially capable of replacing the cognate protein in its interaction with an Ab.

bioinformatics

IsoProt: A fully reproducible one-stop-shop for the analysis of iTRAQ/TMT data

Reproducibility has become a major concern in biomedical research. In proteomics, bioinformatic workflows can quickly consist of multiple software tools each with its own set of parameters. Their usage involves the definition of often hundreds of parameters as well as data operations to ensure tool interoperability. Hence a manuscripts methods section is often insufficient to completely describe and reproduce a data analysis workflow. Here we present IsoProt: A complete and reproducible bioinformatic workflow deployed on a portable container environment to analyse data from isobarically-labeled, quantitative proteomics experiments. The workflow uses only open source tools and provides a user-friendly and interactive browser interface to configure and execute the different operations. Once the workflow is executed, the results including the R code to perform statistical analyses can be downloaded as an HTML or PDF document providing a complete record of the performed analyses. IsoProt therefore represents a reproducible bioinformatics workflow that will yield identical results on any computer platform.

bioinformatics

Epigenetic Regulation of Gene Expression in Cancer: Techniques, Resources, and Analysis

Cancer is a complex disease, driven by aberrant activity in numerous signaling pathways in even individual malignant cells. Epigenetic changes are critical mediators of these functional changes that drive and maintain the malignant phenotype. Changes in DNA methylation, histone acetylation and methylation, non-coding RNAs, post-translational modifications are all epigenetic drivers in cancer, independent of changes in the DNA sequence. These epigenetic alterations, once thought to be crucial only for the malignant phenotype maintenance, are now recognized as critical also for disrupting essential pathways that protect the cells from uncontrolled growth, longer survival and establishment in distant sites from the original tissue. In this review, we focus on DNA methylation and chromatin structure in cancer. While associated with cancer, the precise functional role of these alterations is an area of active research using emerging high-throughput approaches and bioinformatics analysis tools. Therefore, this review describes these high-throughput measurement technologies, public domain databases for high-throughput epigenetic data in tumors and model systems, and bioinformatics algorithms for their analysis. Advances in bioinformatics data integration techniques that combine these epigenetic data with genomics data are essential to infer the function of specific epigenetic alterations in cancer, and are therefore also a focus of this review. Future studies using these emerging technologies will elucidate how alterations in the cancer epigenome cooperate with genetic aberrations to cause tumorigenesis initiation and progression. This deeper understanding is essential to future studies that will precisely infer patients prognosis and select patients who will be responsive to emerging epigenetic therapies.

cancer biology

Novel mutations associated with autosomal dominant congenital cataract are identified in Chinese families

PurposeAs the leading cause of the impairment of vision of children, congenital cataract is considered as a hereditary disease, especially autosomal dominant congenital cataract (ADCC). The purpose of this study is to identify the genetic defect of six Chinese families with ADCC.\n\nSubjects and MethodsSix Chinese families with ADCC were recruited in the study. (103 members in total, 96 members alive, 27 patients in total) Genomic DNA samples extracting from probands peripheral blood cells were captured the mutations using a specific eye disease enrichment panel with next generation sequencing. After initial pathogenicity prediction, sites with specific pathogenicity were screened for further validation. Sanger sequencing was conducted in the other individuals in the families and other 100 normal controls. Mutations definitely related with ADCC will then be analyzed by bioinformatics analysis. The pathogenic effect of the amino acid changes and structural and functional changes of the proteins were finally analyzed by bioinformatics analysis.\n\nResultsSeven mutations in six candidate genes associated with ADCC of six families were detected (MYH9 c.4150G>C, CRYBA4 c.169T>C, RPGRRIP1 c.2669G>A, WFS1 c.1235T>C, CRYBA4 c.26C>T, EPHA2 c.2663+1G>A, and PAX6 c.11-2A>G). All the seven mutations were only detected on affected individuals in the families. Among them there are three novel mutations (MYH9 c.4150G>C, CRYBA4 c.169T>C, RPGRRIP1 c.2669G>A) and four that have been reported (WFS1 c.1235T>C, CRYBA4 c.26C>T, EPHA2 c.2663+1G>A, and PAX6 c.11-2A>G). RPGRIP1 (c.2669G>A) mutation and CRYBA4 (c.26C>T) mutation are predicted to be benign according to bioinformatics analysis while the other five mutations (EPHA2, PAX6, MYH9, CRYBA4 c.169T>C, WFS1) are thought to be pathogenic.\n\nConclusionWe report two novel heterozygous mutations (MYH9 c.4150G>C and CRYBA4 c.169T>C) in six Chinese families supporting their vital roles in causing ADCC.

genomics

Improving Protein Docking with Constraint Programming and Coevolution Data

BackgroundConstraint programming (CP) is usually seen as a rigid approach, focusing on crisp, precise, distinctions between what is allowed as a solution and what is not. At first sight, this makes it seem inadequate for bioinformatics applications that rely mostly on statistical parameters and optimization. The prediction of protein interactions, or protein docking, is one such application. And this apparent problem with CP is particularly evident when constraints are provided by noisy data, as it is the case when using the statistical analysis of Multiple Sequence Alignments (MSA) to extract coevolution information. The goal of this paper is to show that this first impression is misleading and that CP is a useful technique for improving protein docking even with data as vague and noisy as the coevolution indicators that can be inferred from MSA.\n\nResultsHere we focus on the study of two protein complexes. In one case we used a simplified estimator of interaction propensity to infer a set of five candidate residues for the interface and used that set to constrain the docking models. Even with this simplified approach and considering only the interface of one of the partners, there is a visible focusing of the models around the correct configuration. Considering a set of 400 models with the best geometric contacts, this constraint increases the number of models close to the target (RMSD {inverted exclamation}5[A]) from 2 to 5 and decreases the RMSD of all retained models from 26[A] to 17.5[A]. For the other example we used a more standard estimate of coevolving residues, from the Co-Evolution Analysis using Protein Sequences (CAPS) software. Using a group of three residues identified from the sequence alignment as potentially co-evolving to constrain the search, the number of complexes similar to the target among the 50 highest scoring docking models increased from 3 in the unconstrained docking to 30 in the constrained docking.\n\nConclusionsAlthough only a proof-of-concept application, our results show that, with suitably designed constraints, CP allows us to integrate coevolution data, which can be inferred from databases of protein sequences, even though the data is noisy and often \"fuzzy\", with no well-defined discontinuities. This also shows, more generally, that CP in bioinformatics needs not be limited to the more crisp cases of finite domains and explicit rules but can also be applied to a broader range of problems that depend on statistical measurements and continuous data.

Bioinformatics