Search bioRxiv⌕ Search

Biology subjects

Littlefield, S. B.

Publications and source records attributed to Littlefield, S. B..

2 recordsLinked to original sources

Protein Language Models Capture Structural and Functional Epistasis in a Zero-Shot Setting

Protein language models (PLMs) learn from large collections of natural sequences and achieve striking success across prediction tasks, yet it remains unclear what biological principles underlie their representations. We use epistasis, the dependence of a mutations effect on its sequence context, as a lens to probe what PLMs capture about proteins. Comparing PLM-derived scores with deep mutational scanning data, we find that epistasis emerges naturally from pretrained models, without supervision on experimental fitness. Raw model scores align with residue-residue contacts, indicating that PLMs internalize structural proximity. Applying a nonlinear transformation to bring model outputs onto the experimental scale, however, shifts the signal toward functional couplings between distant sites. These findings show that PLMs capture both structural and functional dependencies from sequence data alone, and that epistasis provides a powerful window into the biological principles embedded in their representations.

bioinformatics↗

An unsupervised framework for comparing SARS-CoV-2 protein sequences using LLMs

AO_SCPLOWBSTRACTC_SCPLOWThe severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) pandemic led to 700 million infections and 7 million deaths worldwide. While studying these viruses, scientists developed a large amount of sequencing data that was made available to researchers. Large language models (LLMs) are pre-trained on large databases of proteins and prior work has shown its use in studying the structure and function of proteins. This paper proposes an unsupervised framework for characterizing SARS-CoV-2 sequences using large language models. First, we perform a comparison of several language models previously proposed by other authors. This step is used to determine how clustering and classification approaches perform on SARS-CoV-2 sequence embeddings. In this paper, we focus on surface glycoprotein sequences, also known as spike proteins in SARS-CoV-2 because scientists have previously studied their involvement in being recognized by the human immune system. Our contrastive learning framework is trained in an unsupervised manner, leveraging the Levenshtein distance from pairwise alignment of sequences when the contrastive loss is computed by the Siamese Neural Network. The final part of this paper focuses on a comparison with a previous approach on a test dataset containing data from the latter part of the pandemic. In the prediction of emerging variants, the proposed LLM-based approach shows an improvement of 0.2 in terms of the adjusted rand index clustering compared to a previously proposed approach. This shows the potential of applying large language models to this field.

bioinformatics↗