The Dayhoff Atlas: scaling sequence diversity for improved protein generation
Organized information powers modern biology, a framework pioneered by Margaret Dayhoffs Atlas of Protein Sequence and Structure and advanced by todays databases and computational methods. Here, we extend this paradigm for the AI era, presenting the Dayhoff Atlas of protein sequence data and generative models to accelerate protein biology and design. The Atlas introduces GigaRef, the largest open dataset of natural proteins, spanning 3.34B genomic and metagenomic sequences across 1.70B clusters, and BackboneRef, which distills structural information from 240,811 synthetic backbones into 46M synthetic sequences. Leveraging these datasets, we trained the Dayhoff protein language models, which can predict mutation effects, scaffold structural motifs, and generate novel proteins within families. Training on metagenomic and structure-based synthetic sequences increased the expression rates of generated proteins, demonstrating the value of data diversity and scale. We release the Dayhoff Atlas code, datasets, and models under a permissive license to empower computation in protein design.