Search bioRxiv⌕ Search

Biology subjects

Gracia I Carmona, O.

Publications and source records attributed to Gracia I Carmona, O..

3 recordsLinked to original sources

Assessing structure-function impacts on Vitellogenin by leveraging allelic variant occurring in honey bee subspecies Apis mellifera meliffera

Computational advances involving artificial intelligence (AI) and successful experimental state-of-the-art structure determination can provide detailed pictures of large and complex protein structures, and their variations. A standout case is Vitellogenin (Vg) derived from the honey bee (Apis mellifera). Vg is an essential protein for reproduction in almost all egg-laying animals, and can in addition regulate behavior and provide immunological support in some species, including the honey bee. Information is limited in terms of the structure-function relationships that underlie Vgs pleiotropic functions, at least in part because this protein is not expressed in the best developed gene-editing models, such as fruit flies and mice. However, naturally occurring allelic variation in Vg can provide some insight, i.e., into which changes (mutations) are allowable (present at some frequency) vs. likely not allowable (not present or only at low frequency). Here, we leverage a unique dataset of 1,086 fully sequenced Vg alleles from honey bees in 15 countries. We identify a population-specific 9 nucleotide deletion in a locally endangered honey bee subspecies (A. m. mellifera) that impacts a loop structure in a central Vg domain. Due to the A. m. mellifera population history of near extinction and human intervention, an assessment of this Vg variant is not only theoretically interesting but also relevant for subspecies conservation efforts. Using structural bioinformatics, molecular dynamics simulations, and a transformer-based indel predictor (IndeLLM), we demonstrate that Vg protein structure and stability can be maintained despite the deletion. Our approach also reveals the dynamic nature of specific regions in Vg for the first time. Generalizable results may extend to other egg-laying animals of ecological and economic importance.

bioinformatics↗

Leveraging protein language models and scoring function for Indel characterisation and transfer learning

1.Protein language models (PLMs) are increasingly used to assess the impact of genetic variation on proteins. By leveraging sequence information alone, PLMs achieve high performance and accuracy and can outperform traditional pathogenicity predictors specifically designed to identify harmful variants contributing to diseases. PLMs can perform zero-shot inference, making predictions without task-specific fine-tuning, offering a simpler and less overfitting-prone alternative to complex methods. However, studying in-frame insertions and deletions (indels) with PLMs remains challenging. Indels alter protein length, making direct comparisons between wildtype and mutant sequences not straightforward. Additionally, indel pathogenicity is less studied than other genetic variants, such as single nucleotide variants, resulting in a lack of annotated datasets. Despite these challenges, approaches that leverage PLMs through transfer learning have emerged, making it possible to capture the features needed for more accurate predictions. Still, the current approaches are limited in terms of allowed organisms, indel length, and interpretability. In this work, we devise a new scoring approach for indel pathogenicity prediction (IndeLLM) that provides a solution for the difference in protein lengths. Our method only uses sequence information and zero-shot inference with a fraction of computing time while achieving performances similar to other indel pathogenicity predictors. We used our approach to construct a simple transfer learning approach for a Siamese network, which outperformed all tested indel pathogenicity prediction methods (Matthews correlation coefficient = 0.77). IndeLLM is universally applicable across species since PLMs are trained on diverse protein sequences. To enhance accessibility, we designed a plug-and-play Google Colab notebook that allows easy use of IndeLLM and visualisation of the impact of indels on protein sequence and structure. The tool is available on GitHub https://github.com/OriolGraCar/IndeLLM and Colab https://colab.research.google.com/drive/1CgwprttaNFR_KeJGyFzP0a0C9Y wc4P. Graphical Abstract, if needed, or logo til include on Google Colab O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=159 SRC="FIGDIR/small/642715v1_ufig1.gif" ALT="Figure 1"> View larger version (37K): org.highwire.dtl.DTLVardef@e5d1b0org.highwire.dtl.DTLVardef@299b19org.highwire.dtl.DTLVardef@1859f02org.highwire.dtl.DTLVardef@18a7ba9_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

TITINdb2 - Expanding Annotation and Structural Information for Protein Variants in the Giant Sarcomeric Protein Titin

Structured AbstractO_ST_ABSSummaryC_ST_ABSWe present TITINdb2, an update to the TITINdb database previously constructed to facilitate the identification of pathogenic missense variants in the giant protein titin, which are associated with a variety of skeletal and cardiac myopathies. The database and web portal have been substantially revised and include the following new features: (i) an increase in computational annotation from 4 to 20 variant impact predictors, available through a new custom data table dialogue; (ii) thorough structural coverage of single domains with AlphaFold2 predicted models; (iii) newly predicted domain-domain interface annotations; (iv) an expanded in silico saturation mutagenesis incorporating 4 variant impact predictors; (v) a comprehensive overhaul of available data, including population data sources and variants reported pathogenic in the literature; (vi) A curated mapping of existing protein, transcript and chromosomal sequence positions and a new variant conversion tool to translate variants in one format to any other format. Availability and ImplementationDatabase accessible via titindb.kcl.ac.uk/TITINdb/ ContactFranca Fraternali (f.fraternali@ucl.ac.uk) Supplementary InformationAvailable

bioinformatics↗