Search bioRxivSearch

Biology subjects

Basile, W.

Publications and source records attributed to Basile, W..

3 recordsLinked to original sources

Why do eukaryotic proteins contain more intrinsically disordered regions?

Intrinsic disorder is much more abundant in eukaryotic than in prokaryotic proteins. However, the reason behind this is unclear. It has been proposed that the disordered regions are functionally important for regulation in eukaryotes, but it has also been proposed that the difference is a result of lower selective pressure in eukaryotes. Almost all studies intrinsic disorder is predicted from the amino acid sequence of a protein. Therefore, there should exist an underlying difference in the amino acid distributions between eukaryotic and prokaryotic proteins causing the predicted difference in intrinsic disorder. To obtain a better understanding of why eukaryotic proteins contain more intrinsically disordered regions we compare proteins from complete eukaryotic and prokaryotic proteomes.\n\nHere, we show that the difference in intrinsic disorder origin from differences in the linker regions. Eukaryotic proteins have more extended linker regions and, in particular, the eukaryotic linker regions are more disordered. The average eukaryotic protein is about 500 residues long; it contains 250 residues in linker regions, of which 80 are disordered. In comparison, prokaryotic proteins are about 350 residues long and only have 100-110 residues in linker regions, and less than 10 of these are intrinsically disordered.\n\nFurther, we show that there is no systematic increase in the frequency of disorder-promoting residues in eukaryotic linker regions. Instead, the difference in frequency of only three amino acids seems to lie behind the difference. The most significant difference is that eukaryotic linkers contain about 9% serine, while prokaryotic linkers have roughly 6.5%. Eukaryotic linkers also contain about 2% more proline and 2-3% fewer isoleucine residues. The reason why primarily these amino acids vary in frequency is not apparent, but it cannot be excluded that the difference is serine is related to the increased need for regulation through phosphorylation and that the proline difference is related to increase of eukaryotic specific repeats.

bioinformatics

The number of orphans in yeast and fly is drastically reduced by using combining searches in both proteomes and genomes

The identification of de novo created genes is important as it provides a glimpse on the evolutionary processes of gene creation. Potential de novo created genes are identified by selecting genes that have no homologs outside a particular species, but for an accurate detection this identification needs to be correct.\n\nGenes without any homologs are often referred to as orphans; in addition to de novo created ones, fast evolving genes or genes lost in all related genomes might also be classified as orphans. The identification of orphans is dependent on: (i) a method to detect homologs and (ii) a database including genes from related genomes.\n\nHere, we set out to investigate how the detection of orphans is influenced by these two factors. Using Saccharomyces cerevisiae we identify that best strategy is to use a combination of searching annotated proteins and a six-frame translation of all ORFs from closely related genomes. Using this strategy we obtain a set of 54 orphans in Drosophila melanogaster and 38 in Drosophila pseudoobscura, significantly less than what is reported in some earlier studies.

bioinformatics

High GC Content Causes De Novo Created Proteins to be Intrinsically Disordered

De novo creation of protein coding genes involves the formation of short ORFs from noncoding regions; some of these ORFs might then become fixed in the populationThese orphan proteins need to, at the bare minimum, not cause serious harm to the organism, meaning that they should for instance not aggregate. Therefore, although the creation of short ORFs could be truly random, the fixation should be subjected to some selective pressure. The selective forces acting on orphan proteins have been elusive, and contradictory results have been reported. In Drosophila young proteins are more disordered than ancient ones, while the opposite trend is present in yeast. To the best of our knowledge no valid explanation for this difference has been proposed.\n\nTo solve this riddle we studied structural properties and age of proteins in 187 eukaryotic organisms. We find that, with the exception of length, there are only small differences in the properties between proteins of different ages. However, when we take the GC content into account we noted that it could explain the opposite trends observed for orphans in yeast (low GC) and Drosophila (high GC). GC content is correlated with codons coding for disorder promoting amino acids. This leads us to propose that intrinsic disorder is not a strong determining factor for fixation of orphan proteins. Instead these proteins largely resemble random proteins given a particular GC level. During evolution the properties of a protein change faster than the GC level causing the relationship between disorder and GC to gradually weaken.\n\nAuthor SummaryWe show that the GC content of a genome is of great importance for the properties of an orphan protein. GC content affects the frequency of the codons and this affects the probability for each amino acid to be included in a de novo created protein. The codons encoding for Ala, Pro and Gly contain 80% GC, while codons for Lys, Phe, Asn, Tyr and Ile contain 20% or less. The three high GC amino acids are all disorder promoting, while Phe, Tyr and Ile are order promoting. Therefore, random protein sequences at a high GC will be more disordered than the ones created at a low GC. The structural properties of the youngest proteins match to a large degree the properties of random proteins when the GC content is taken into account. In contrast, structural properties of ancient proteins only show a weak correlation with GC content. This suggests that even after fixation in the population, proteins largely resemble random proteins given a certain GC content. Thereafter, during evolution the correlation between structural properties and GC weakens.

evolutionary biology