Search bioRxiv⌕ Search

Biology subjects

Gangavarapu, P.

Publications and source records attributed to Gangavarapu, P..

2 recordsLinked to original sources

The emergence and molecular evolution of H5N1 influenza viruses in United States dairy cattle

Prior to 2024, highly pathogenic avian influenza H5N1 clade 2.3.4.4b viruses circulated predominantly in wild birds and poultry. In 2024 and 2025, 2.3.4.4b genotypes B3.13 and D1.1 were detected in United States dairy cattle. Using whole-genome and segment-specific phylodynamic inference, we estimate that B3.13 and D1.1 spilled over from wild birds into dairy cattle in late 2023 and late 2024, respectively. Spillover occurred shortly after the formation of the reassortant genotypes and was followed by months of cryptic transmission prior to detection. We found that both B3.13 and D1.1 evolved at higher rates in cattle relative to birds, primarily due to relaxed purifying selection. Site-specific analyses identified genomic sites under positive selection in cattle relative to birds, indicating adaptation and likely contributing to improved viral fitness after spillover. Intensified genomic surveillance in dairy cattle is essential as population immunity introduces additional selection pressures, with ever-changing risk for human emergence.

evolutionary biology↗

Data-optimal scaling of paired antibody language models

Scaling laws for large language models in natural language domains are typically derived under the assumption that performance is primarily compute-constrained. In contrast, antibody language models (AbLMs) trained on paired sequences are primarily data-limited, thus requiring different considerations. To explore how model size and data scale affect AbLM performance, we trained 15 AbLMs across all pairwise combinations of five model sizes and three training data sizes. From these experiments, we derive an AbLM-specific scaling law and estimate that training a data-optimal AbLM equivalent of the highly performant 650M-parameter ESM-2 protein language model would require [~]5.5 million paired antibody sequences. Evaluation on multiple downstream classification tasks revealed that significant performance gains emerged only with sufficiently large model size, suggesting that in data-limited domains, improved performance depends jointly on both model scale and data volume.

immunology↗