bioRxiv · 10.1101/2024.09.23.614632
Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data
Abstract
Genomic sequence-to-expression deep learning models, which are trained to predict gene expression and other molecular phenotypes across the reference genome, have recently been shown to have poor out-of-the-box performance in predicting gene expression variation across individuals based on their personal genome sequences. Here we explore whether additional training (fine-tuning) on paired personal genome and transcriptome data improves the performance of such sequence-to-expression models. Using Enformer as a representative pretrained model, we explore various fine-tuning strategies. Our results show that fine-tuning improves cross-individual prediction performance over the baseline Enformer model for held-out individuals on genes seen during fine-tuning, with comparable performance to variant-based linear models commonly used in transcriptome-wide association studies. However, fine-tuning does not improve model generalizability on held-out genes, which contain sequences and variants unseen during fine-tuning, highlighting a remaining open challenge in the field.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Rastogi, R., Reddy, A. J., Chung, R., Ioannidis, N. M.. 2024-09-25. Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data. https://doi.org/10.1101/2024.09.23.614632
Cite the original work for its findings. Save a collection to share your selection of sources.