Search bioRxivSearch

Biology subjects

Di Lena, P.

Publications and source records attributed to Di Lena, P..

2 recordsLinked to original sources

Fold recognition by scoring protein map similarities using the congruence coefficient

MotivationProtein fold recognition is a key step for template-based modeling approaches to protein structure prediction. Although closely related folds can be easily identified by sequence homology search in sequence databases, fold recognition is notoriously more difficult when it involves the identification of distantly related homologues. Recent progress in residue-residue contact and distance prediction opens up the possibility of improving fold recognition by using structural information contained in predicted distance and contact maps. ResultsHere we propose to use the congruence coefficient as a metric of similarity between maps. We prove that this metric has several interesting mathematical properties which allow one to compute in polynomial time its exact mean and variance over all possible (exponentially many) alignments between two symmetric matrices, and assess the statistical significance of similarity between aligned maps. We perform fold recognition tests by recovering predicted target contact/distance maps from the two most recent CASP editions and over 27,000 non-homologous structural templates from the ECOD database. On this large benchmark, we compare fold recognition performances of different alignment tools with their own similarity scores against those obtained using the congruence coefficient. We show that the congruence coefficient overall improves fold recognition over other methods, proving its effectiveness as a general similarity metric for protein map comparison. AvailabilityThe software CCpro is available as part of the Scratch suite http://scratch.proteomics.ics.uci.edu/

bioinformatics

NETGE-PLUS: standard and network-based gene enrichment analysis in human and model organisms

Omics techniques provide a spectrum of information that needs to be disentangled to characterize complex traits at the molecular level. The gap between genotype and phenotype must be closed by reconciling the genome information with the set of molecular pathways and biological processes describing the phenotype. In dealing with this problem, gene enrichment analysis has become the most widely adopted strategy. Here, we present NETGE-PLUS, a web-server for standard and network-based functional interpretation of gene sets of human and of model organisms, including S. scrofa, S. cerevisiae, E. coli and A. thaliana. NETGE-PLUS enables the functional enrichment of both simple and ranked lists of genes, also introducing the possibility of exploring relationships among KEGG pathways. A web interface makes data retrieval complete and user-friendly. NETGE-PLUS is publicly available at http://net-ge2.biocomp.unibo.it

bioinformatics