Search bioRxiv⌕ Search

Biology subjects

Thessen, A.

Publications and source records attributed to Thessen, A..

2 recordsLinked to original sources

Environment-Aware DNA Language Model for Stress-Responsive Genomic Prioritization in Maize

Abiotic stresses such as heat and drought severely reduce maize productivity, yet identifying genomic regions that confer stress resilience remains a challenge. Inspired by advances in Large Language Models (LLMs), Genomic Foundation Models (GFMs) have recently emerged as a promising approach for capturing regulatory patterns through large-scale pre-training on DNA sequences. However, their application to plant stress-response analysis remains unexplored. This study presents an environment-aware DNA-LLM that adapts AgroNT, a transformer-based GFM pre-trained on diverse plant genomes, by incorporating stress-specific prompt tokens. Through parameter-efficient fine-tuning, the model learns stress-conditioned sequence representations that form distinct clusters in the embedding space across environmental contexts. By combining stress-induced shifts in these sequence representations relative to control conditions with transformer attention patterns, we prioritized putative heat- and drought-responsive genomic regions associated with grain yield in the Genomes-to-Fields (G2F) panel. Prioritized regions were supported by spatiotemporal differential gene-expression evidence and overlap with stress-associated quantitative trait loci. They were further characterized through transcription-factor family analysis and regulatory motif enrichment. Attention-guided analysis additionally identified stress-associated motifs enriched within model-emphasized sequence regions. Overall, the prioritized loci were proximal to genes involved in transcriptional regulation, signaling, and metabolic pathways relevant to abiotic-stress adaptation, demonstrating the potential of stress-conditioned transformer-based sequence modeling for environment-aware genome-to-phenome analysis.

genomics↗

Post-GWAS Prioritization of Genome-Phenome Associations in Sorghum

Genome-Wide Association Studies (GWAS) are widely used to infer the genetic basis of traits in organisms, yet selecting appropriate thresholds for analysis remains a significant challenge. In this study, we developed the Sequential SNP Prioritization Algorithm (SSPA) to elucidate the genetic underpinnings of two key phenotypes in Sorghum bicolor: maximum canopy height and maximum growth rate. Utilizing a subset of the Sorghum Bioenergy Association Panel cultivated at the Maricopa Agricultural Center in Arizona, our objective was to employ GWAS with specific permissive-filtered thresholds to identify the genetic markers associated with these traits, allowing for a broader collection of explanatory candidate genes. Following this, our proposed method incorporates a feature engineering approach based on statistical correlation coefficient to reveal patterns between phenotypic similarity and genetic proximity across 274 accessions. This approach helps prioritize Single Nucleotide Polymorphisms (SNPs) likely to be associated with the studied phenotype. Additionally, we evaluated the impact of SSPA by considering all variants (SNPs) as inputs, without any GWAS filtering, as a complementary analysis. Empirical evidence including ontology- based gene function, spatial and temporal expression, and similarity to known homologs, demonstrated that SSPA effectively prioritizes SNPs and genes influencing the phenotype of interest, providing valuable insights for functional genetics research.

genomics↗