Search bioRxiv⌕ Search

Biology subjects

Khaiwal, S.

Publications and source records attributed to Khaiwal, S..

2 recordsLinked to original sources

The adaptive molecular landscape of reprogrammed telomeric sequences

Telomeric sequences vary across the tree of life and intimately co-evolve with telomere-binding protein complexes. However, the molecular mechanisms allowing organisms to adapt to new telomeric sequences are difficult to gauge from extant species. Here, we reprogrammed multiple yeast lines to human-like telomeric repeats to unveil their molecular and fitness response to novel telomeres. Initially, the exchange of telomere sequences resulted in genome instability, proteome remodelling and severe fitness decline. However, adaptive evolution experiments selected for repeated mutations that drove adaptation to the humanized telomeres. These consisted of the recurrent amplification of the telomere-binding protein TBF1, by complex aneuploidies, or in repeated mutations that attenuate the DNA damage response. Overall, our results outline a response that defines the adaptive molecular landscape to novel telomeric sequences.

evolutionary biology↗

Predicting the natural yeast phenotypic landscape with machine learning

Most organisms traits result from the complex interplay of many genetic and environmental factors, making their prediction from genotypes difficult. Here, we used machine learning models to explore genotype-phenotype connections for 223 life history traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains. Firstly, we used genome-wide association studies to connect genetic variants with the phenotypes. Next, we benchmarked an automated machine learning pipeline that includes preprocessing, feature selection, and hyperparameters optimization in combination with multiple linear and complex machine learning methods. We determined gradient boosting machines as best performing in 65% of predictions and pangenome as best predictor, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The accuracy broadly varied among the phenotypes (r = 0.2-0.9), consistent with varying levels of complexity, with stress resistance being easier to predict compared to growth across carbon and nitrogen nutrients. While no specific genomic features could be linked to the predictions for most phenotypes, machine learning identifies high-impact variants with established relationships to phenotypes despite being rare in the population. Near-perfect accuracies (r>0.95) were achieved when other phenomics data were used to aid predictions, suggesting shared useful information can be conveyed across phenotypes. Overall, our study underscores the power of machine learning to interpret the functional outcome of genetic variants.

genetics↗