Search bioRxivSearch

Biology subjects

Nielly-Thibault, L.

Publications and source records attributed to Nielly-Thibault, L..

3 recordsLinked to original sources

The high turnover of ribosome-associated transcripts from de novo ORFs produces gene-like characteristics available for de novo gene emergence in wild yeast populations

Little is known about the rate of emergence of genes de novo, how they spread in populations and what their initial properties are. We examined wild yeast (Saccharomyces paradoxus) populations to characterize the diversity and turnover of intergenic ORFs over short evolutionary time-scales. With ~34,000 intergenic ORFs per individual genome for a total of ~64,000 orthogroups identified, we found de novo ORF formation to have a lower estimated turnover rate than gene duplication. Hundreds of intergenic ORFs show translation signatures similar to canonical genes. However, they have lower translation efficiency, which could reflect a mechanism to reduce their production cost or simply a lack of optimization. We experimentally confirmed the translation of many of these ORFs in laboratory conditions using a reporter assay. Translated intergenic ORFs tend to display low expression levels with sequence properties that generally are close to expectations based on intergenic sequences. However, some of the very recent translated intergenic ORFs, which appeared less than 110 Kya ago, already show gene- like characteristics, suggesting that the raw material for functional innovations could appear over short evolutionary time-scales.

evolutionary biology

Differences between the de novo proteome and its non-functional precursor can result from neutral constraints on its birth process, not necessarily from natural selection alone

Proteins are among the most important constituents of biological systems. Because all proteins ultimately evolved from previously non-coding DNA, the properties of these non-coding sequences and how they shape the birth of novel proteins are also expected to influence the organization of biological networks. When trying to explain and predict the properties of novel proteins, it is of particular importance to distinguish the contributions of natural selection and other evolutionary forces. Studies in the field typically use non-coding DNA and GC-content-based random-sequence models to generate random expectations for the properties of novel functional proteins. Deviations from these expectations have been interpreted as the result of natural selection. However, interpreting such deviations requires a yet-unattained understanding of the raw material of de novo gene birth and its relation to novel functional proteins. We mathematically show how the importance of the \"junk\" polypeptides that make up this raw material goes beyond their average properties and their filtering by natural selection. We find that the mean of any property among novel functional proteins also depends on its variance among junk polypeptides and its correlation with their rate of evolutionary turnover. In order to exemplify the use of our general theoretical results, we combine them with a simple model that predicts the means and variances of the properties of junk polypeptides from the genomic GC content alone. Under this model, we predict the effect of GC content on the mean length and mean intrinsic disorder of novel functional proteins as a function of evolutionary parameters. We use these predictions to formulate new evolutionary interpretations of published data on the length and intrinsic disorder of novel functional proteins. This work provides a theoretical framework that can serve as a guide for the prediction and interpretation of past and future results in the study of novel proteins and their properties under various evolutionary models. Our results provide the foundation for a better understanding of the properties of cellular networks through the evolutionary origin of their components.

evolutionary biology

Highly efficient CRISPR gene editing in yeast enabled by double selection

CRISPR-Cas9 loss of function (LOF) and base editing screens are powerful tools in genetics and genomics. Yeast is one of the main models in genetics and genomics, yet large-scale approaches remain to be developed in this species because of low mutagenesis rates without donor DNA. We developed a double selection strategy based on co-selection that increases LOF mutation rates, both for CRISPR-Cas9 and the Target-AID base editor. We constructed the pDYSCKO vector, which is amenable to high throughput double selection for both approaches. Using modeling, we show that this improvement provides the required increased in detection power to measure the fitness effects of thousands of mutations in typical yeast pooled screens. We also show that multiplex genome editing with Cas9 causes programmable chromosomal translocations at high frequency, suggesting that multiplex editing should be performed with caution and that base-editors could be preferable tools for LOF screens.

synthetic biology