Search bioRxiv⌕ Search

Biology subjects

Krähenbühl, P.

Publications and source records attributed to Krähenbühl, P..

2 recordsLinked to original sources

Distilling structural representations into protein sequence models

Protein language models, like the popular ESM2, are widely used tools for extracting evolution-based protein representations and have achieved significant success on downstream biological tasks. Representations based on sequence and structure models, however, show significant performance differences depending on the downstream task. A major open problem is to obtain representations that best capture both the evolutionary and structural properties of proteins in general. Here we introduce Implicit Structure Model (ISM), a sequence-only input model with structurally-enriched representations that outperforms state-of-the-art sequence models on several well-studied benchmarks including mutation stability assessment and structure prediction. Our key innovations are a microenvironment-based autoencoder for generating structure tokens and a self-supervised training objective that distills these tokens into ESM2s pre-trained model. We have made ISMs structure-enriched weights easily available: integrating ISM into any application using ESM2 requires changing only a single line of code. Our code is available at https://github.com/jozhang97/ISM.

bioinformatics↗

Image-based phenomic prediction can provide valuable decision support in wheat breeding

Traditionally, breeders selection decisions in early generations are largely based on visual observations in the field. With the advent of affordable genome sequencing and high-throughput phenotyping technologies, enhancing breeders ratings with such information became attractive. In this research, it is hypothesized that GxE interactions of secondary traits (i.e., growth dynamics traits) are less complex than those of related target traits (e.g., yield). Thus, phenomic selection (PS) may allow selecting for genotypes with beneficial response-pattern in a defined population of environments. A set of 45 winter wheat varieties was grown at five year-sites and analyzed with linear and factor-analytic (FA) mixed models to estimate GxE interactions of secondary and target traits. The dynamic development of drone-derived plant height, leaf area and tiller density estimations was used to estimate the timing of key stages, quantities at defined time points, and temperature dose-response curve parameters. Most of these secondary traits and grain protein content showed little GxE interactions. In contrast, the modeling of GxE for yield required a FA model with two factors. A trained PS model predicted overall yield performance, yield stability and grain protein content with correlations of 0.43, 0.30 and 0.34. While these accuracies are modest and do not outperform well-trained GS models, PS additionally provided insights into the physiological basis of target traits. An ideotype was identified that potentially avoids the negative pleiotropic effects between yield and protein content. Key messageGenotype-by-environment interactions of secondary traits based on high-throughput field phenotyping are less complex than those of target traits, allowing for a phenomic selection in unreplicated early generation trials.

physiology↗