Search bioRxivSearch

Biology subjects

Elofsson, A.

Publications and source records attributed to Elofsson, A..

7 recordsLinked to original sources

PconsC4: fast, free, easy, and accurate contact predictions.

MotivationResidue contact prediction was revolutionized recently by the introduction of direct coupling analysis (DCA). Further improvements, in particular for small families, have been obtained by the combination of DCA and deep learning methods. However, existing deep learning contact prediction methods often rely on a number of external programs and are therefore computationally expensive.\n\nResultsHere, we introduce a novel contact predictor, PconsC4, which performs on par with state of the art methods. PconsC4 is heavily optimized, does not use any external programs and therefore is significantly faster and easier to use than other methods.\n\nAvailabilityPconsC4 is freely available under the GPL license from https://github.com/ElofssonLab/PconsC4. Installation is easy using the pip command and works on any system with Python 3.5 or later and a modern GCC compiler.\n\nContactarne@bioinfo.se

bioinformatics

Improved protein model quality assessments by changing the target function

Protein modeling quality is an important part of protein structure prediction. We have for more than a decade developed a set of methods for this problem. We have used various types of description of the protein and different machine learning methodologies. However, common to all these methods has been the target function used for training. The target function in ProQ describes the local quality of a residue in a protein model. In all versions of ProQ the target function has been the S-score. However, other quality estimation functions also exist, which can be divided into superposition- and contact-based methods. The superposition-based methods, such as S-score, are based on a rigid body superposition of a protein model and the native structure, while the contact-based methods compare the local environment of each residue. Here, we examine the effects of retraining our latest predictor, ProQ3D, using identical inputs but different target functions. We find that the c ntact-based methods are easier to predict and that predictors trained on these measures provide some advantages when it comes to identifying the best model. One possible reason for this is that contact based methods are better at estimating the quality of multi-domain targets. However, training on the S-score gives the best correlation with the GDT_TS score, which is commonly used in CASP to score the global model quality. To take the advantage of both of these features we provide an updated version of ProQ3D that predicts local and global model quality estimates based on different quality estimates.

bioinformatics

Why do eukaryotic proteins contain more intrinsically disordered regions?

Intrinsic disorder is much more abundant in eukaryotic than in prokaryotic proteins. However, the reason behind this is unclear. It has been proposed that the disordered regions are functionally important for regulation in eukaryotes, but it has also been proposed that the difference is a result of lower selective pressure in eukaryotes. Almost all studies intrinsic disorder is predicted from the amino acid sequence of a protein. Therefore, there should exist an underlying difference in the amino acid distributions between eukaryotic and prokaryotic proteins causing the predicted difference in intrinsic disorder. To obtain a better understanding of why eukaryotic proteins contain more intrinsically disordered regions we compare proteins from complete eukaryotic and prokaryotic proteomes.\n\nHere, we show that the difference in intrinsic disorder origin from differences in the linker regions. Eukaryotic proteins have more extended linker regions and, in particular, the eukaryotic linker regions are more disordered. The average eukaryotic protein is about 500 residues long; it contains 250 residues in linker regions, of which 80 are disordered. In comparison, prokaryotic proteins are about 350 residues long and only have 100-110 residues in linker regions, and less than 10 of these are intrinsically disordered.\n\nFurther, we show that there is no systematic increase in the frequency of disorder-promoting residues in eukaryotic linker regions. Instead, the difference in frequency of only three amino acids seems to lie behind the difference. The most significant difference is that eukaryotic linkers contain about 9% serine, while prokaryotic linkers have roughly 6.5%. Eukaryotic linkers also contain about 2% more proline and 2-3% fewer isoleucine residues. The reason why primarily these amino acids vary in frequency is not apparent, but it cannot be excluded that the difference is serine is related to the increased need for regulation through phosphorylation and that the proline difference is related to increase of eukaryotic specific repeats.

bioinformatics

The number of orphans in yeast and fly is drastically reduced by using combining searches in both proteomes and genomes

The identification of de novo created genes is important as it provides a glimpse on the evolutionary processes of gene creation. Potential de novo created genes are identified by selecting genes that have no homologs outside a particular species, but for an accurate detection this identification needs to be correct.\n\nGenes without any homologs are often referred to as orphans; in addition to de novo created ones, fast evolving genes or genes lost in all related genomes might also be classified as orphans. The identification of orphans is dependent on: (i) a method to detect homologs and (ii) a database including genes from related genomes.\n\nHere, we set out to investigate how the detection of orphans is influenced by these two factors. Using Saccharomyces cerevisiae we identify that best strategy is to use a combination of searching annotated proteins and a six-frame translation of all ORFs from closely related genomes. Using this strategy we obtain a set of 54 orphans in Drosophila melanogaster and 38 in Drosophila pseudoobscura, significantly less than what is reported in some earlier studies.

bioinformatics

Methods For Estimation Of Model Accuracy In CASP12

Methods for reliably estimating the quality of 3D models of proteins are essential drivers for the wide adoption and serious acceptance of protein structure predictions by life scientists. In this paper, the most successful groups in CASP12 describe their latest methods for Estimates of Model Accuracy (EMA). We show that pure single model accuracy estimation methods has shown clear progress since CASP11; the three top methods (MESHI, ProQ3, SVMQA) all perform better than the top method of CASP11 (ProQ2). The pure single model accuracy estimation methods outperform quasi-single (ModFOLD6 variations) and consensus methods (Pcons, ModFOLDclust2, Pcomb-domain and Wallner) in model selection, but are still not as good as those methods in absolute model quality estimation and predictions of local quality. Finally, we show that when using contact based model quality measures (CAD, 1DDT) the single model quality methods perform relatively better.

bioinformatics

Large-Scale Structure Prediction By Improved Contact Predictions And Model Quality Assessment

MotivationAccurate contact predictions can be used for predicting the structure of proteins. Until recently these methods were limited to very big protein families, decreasing their utility. However, recent progress by combining direct coupling analysis with machine learning methods has made it possible to predict accurate contact maps for smaller families. To what extent these predictions can be used to produce accurate models of the families is not known.\n\nResultsWe present the PconsFold2 pipeline that uses contact predictions from PconsC3, the CONFOLD folding algorithm and model quality estimations to predict the structure of a protein. We show that the model quality estimation significantly increases the number of models that reliably can be identified. Finally, we apply PconsFold2 to 6379 Pfam families of unknown structure and find that PconsFold2 can, with an estimated 90% specificity, predict the structure of up to 558 Pfam families of unknown structure. Out of these 415 have not been reported before.\n\nAvailabilityDatasets as well as models of all the 558 Pfam families are available at http://c3.pcons.net/. All programs used here are freely available.\n\nContactarne@bioinfo.se\n\nSupplementary informationNo supplementary data

bioinformatics

High GC Content Causes De Novo Created Proteins to be Intrinsically Disordered

De novo creation of protein coding genes involves the formation of short ORFs from noncoding regions; some of these ORFs might then become fixed in the populationThese orphan proteins need to, at the bare minimum, not cause serious harm to the organism, meaning that they should for instance not aggregate. Therefore, although the creation of short ORFs could be truly random, the fixation should be subjected to some selective pressure. The selective forces acting on orphan proteins have been elusive, and contradictory results have been reported. In Drosophila young proteins are more disordered than ancient ones, while the opposite trend is present in yeast. To the best of our knowledge no valid explanation for this difference has been proposed.\n\nTo solve this riddle we studied structural properties and age of proteins in 187 eukaryotic organisms. We find that, with the exception of length, there are only small differences in the properties between proteins of different ages. However, when we take the GC content into account we noted that it could explain the opposite trends observed for orphans in yeast (low GC) and Drosophila (high GC). GC content is correlated with codons coding for disorder promoting amino acids. This leads us to propose that intrinsic disorder is not a strong determining factor for fixation of orphan proteins. Instead these proteins largely resemble random proteins given a particular GC level. During evolution the properties of a protein change faster than the GC level causing the relationship between disorder and GC to gradually weaken.\n\nAuthor SummaryWe show that the GC content of a genome is of great importance for the properties of an orphan protein. GC content affects the frequency of the codons and this affects the probability for each amino acid to be included in a de novo created protein. The codons encoding for Ala, Pro and Gly contain 80% GC, while codons for Lys, Phe, Asn, Tyr and Ile contain 20% or less. The three high GC amino acids are all disorder promoting, while Phe, Tyr and Ile are order promoting. Therefore, random protein sequences at a high GC will be more disordered than the ones created at a low GC. The structural properties of the youngest proteins match to a large degree the properties of random proteins when the GC content is taken into account. In contrast, structural properties of ancient proteins only show a weak correlation with GC content. This suggests that even after fixation in the population, proteins largely resemble random proteins given a certain GC content. Thereafter, during evolution the correlation between structural properties and GC weakens.

evolutionary biology