dnaSORA - A Unified Diffusion Transformer for DNA point clouds
The relatively obscure Hawaiian experiment collapses diverse phenotypes, including nearly all human genetic diseases to a singular Gaussian-like point cloud feature, structuring unstructured information. The uniformity of the feature space provides a straightforward way for AI models to learn all three billion tokens for reading the human genome as a first language. We propose a diffusion transformer, dnaSORA, for learning these features. dnaSORA has generative capacity similar to Stable Diffusion but for DNA point clouds. The models architecture is novel because it is unified; thus, it also functions as a discriminator that uses a frozen latent representation for classification. dnaSORA transfer learns from synthetic data emulating real genome point clouds to classify misrepresented tokens in C. elegans Hawaiian data at state-of-the-art 0.3 Mb resolution. Pre-training large genome models typically requires expensive and difficult-to-obtain genomes. However, our solution provides nearly unlimited synthetic training data at negligible compute costs. Inference for new token assignments (e.g., new diseases) requires genomes from several dozen rather than thousands of individuals. These efficiencies, combined with state-of-the-art resolution, provide a pathway for rapid, massive scaling of token annotation of the entire human genome at orders of magnitude below expected costs.