Search bioRxiv⌕ Search

Biology subjects

Gilliot, P.-A.

Publications and source records attributed to Gilliot, P.-A..

2 recordsLinked to original sources

Transfer learning for cross-context prediction of protein expression from 5'UTR sequence

Model-guided DNA sequence design can accelerate the reprogramming of living cells. It allows us to engineer more complex biological systems by removing the need to physically assemble and test each potential design. While mechanistic models of gene expression have seen some success in supporting this goal, data-centric, deep learning-based approaches often provide more accurate predictions. This accuracy, however, comes at a cost -- a lack of generalisation across genetic and experimental contexts, which has limited their wider use outside the context in which they were trained. Here, we address this issue by demonstrating how a simple transfer learning procedure can effectively tune a pre-trained deep learning model to predict protein translation rate from 5 untranslated region sequence (5UTR) for diverse contexts in Escherichia coli using a small number of new measurements. This allows for important model features learnt from expensive massively parallel reporter assays to be easily transferred to new settings. By releasing our trained deep learning model and complementary calibration procedure, this study acts as a starting point for continually refined model-based sequence design that builds on previous knowledge and future experimental efforts.

synthetic biology↗

Effective design and inference for cell sorting and sequencing based massively parallel reporter assays

The ability to measure the phenotype of millions of different genetic designs using Massively Parallel Reporter Assays (MPRAs) has revolutionised our understanding of genotype-to-phenotype relationships and opened avenues for data-centric approaches to biological design. However, our knowledge of how best to design these costly experiments and the effect that our choices have on the quality of the data produced is lacking. Here, we tackle this issue by developing FORE-CAST, a Python package that supports the accurate simulation of cell-sorting and sequencing based MPRAs and robust maximum like-lihood based inference of genetic design function from MPRA data. We use FORECASTs capabilities to reveal rules for MPRA experimental design that help ensure accurate genotype-to-phenotype links and show how the simulation of MPRA experiments can help us better understand the limits of prediction accuracy when this data is used for training deep learning based classifiers. As the scale and scope of MPRAs grows, tools like FORECAST will help ensure we make informed decisions during their development and the most of the data produced.

synthetic biology↗