bioRxiv · 10.1101/2024.09.18.612131
Pangenome-Informed Language Models for Privacy-Preserving Synthetic Genome Sequence Generation
Abstract
Language Models (LM) have been extensively utilized for learning DNA sequence patterns and generating synthetic sequences. In this paper, we present a novel approach for the generation of synthetic DNA data using pangenomes in combination with LM. We introduce three innovative pangenome-based tokenization schemes that enhance DNA sequence generation. Our experimental results demonstrate the superiority of pangenome-based tokenization over classical methods in generating high-utility synthetic DNA sequences, highlighting significant improvements in training efficiency and sequence quality.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Huang, P., Charton, F., Schmelzle, J.-N. M., Darnell, S. S., Prins, P., Garrison, E., Suh, G. E.. 2024-09-20. Pangenome-Informed Language Models for Privacy-Preserving Synthetic Genome Sequence Generation. https://doi.org/10.1101/2024.09.18.612131
Cite the original work for its findings. Save a collection to share your selection of sources.