Search bioRxiv⌕ Search

Biology subjects

Ansuini, A.

Publications and source records attributed to Ansuini, A..

5 recordsLinked to original sources

A Bayesian framework to infer and cluster mutational signatures leveraging prior biological knowledge.

Mutational signatures provide key insights into cancer mutational processes, but the availability of signature catalogues generated by different groups using distinct methodologies underscores a need for standardisation. We introduce a Bayesian framework that offers a systematic approach to expanding existing signature catalogues for any type of mutational signature, while grouping patients based on shared signature patterns. We demonstrate that this approach can identify both known and novel molecular subtypes across nearly 8,000 samples spanning six cancer types, and show that stratifications derived from signature yield prognostic groups, further enhancing the translational potential of mutational signatures.

cancer biology↗

Unsupervised domain classification of AlphaFold2-predicted protein structures

AO_SCPLOWBSTRACTC_SCPLOWThe release of the AlphaFold database, which contains 214 million predicted protein structures, represents a major leap forward for proteomics and its applications. However, lack of comprehensive protein annotation limits its accessibility and usability. Here, we present DPCstruct, an unsupervised clustering algorithm designed to provide domain-level classification of protein structures. Using structural predictions from AlphaFold2 and comprehensive all-against-all local alignments from Foldseek, DPCstruct identifies and groups recurrent structural motifs into domain clusters. When applied to the Foldseek Cluster database, a representative set of proteins from the AlphaFoldDB, DPCstruct successfully recovers the majority of protein folds catalogued in established databases such as SCOP and CATH. Out of the 28,246 clusters identified by DPCstruct, 24% have no structural or sequence similarity to known protein families. Supported by a modular and efficient implementation, classifying 15 million entries in less than 48 hours, DPCstruct is well suited for large-scale proteomics and metagenomics applications. It also facilitates the rapid incorporation of updates from the latest structural prediction tools, ensuring that the classification remains up-to-date. The DPCstruct pipeline and associated database are freely available in a dedicated repository, enhancing the navigation of the AlphaFoldDB through domain annotations and enabling rapid classification of other protein datasets.

bioinformatics↗

Enhancing predictions of protein stability changes induced by single mutations using MSA-based Language Models

Protein Language Models offer a new perspective for addressing challenges in structural biology, while relying solely on sequence information. Recent studies have investigated their effectiveness in forecasting shifts in thermodynamic stability caused by single amino acid mutations, a task known for its complexity due to the sparse availability of data, constrained by experimental limitations. To tackle this problem, we introduce two key novelties: leveraging a Protein Language Model that incorporates Multiple Sequence Alignments to capture evolutionary information, and using a recently released mega-scale dataset with rigorous data pre-processing to mitigate overfitting. We ensure comprehensive comparisons by fine-tuning various pre-trained models, taking advantage of analyses such as ablation studies and baselines evaluation. Our methodology introduces a stringent policy to reduce the widespread issue of data leakage, rigorously removing sequences from the training set when they exhibit significant similarity with the test set. The MSA Transformer emerges as the most accurate among the models under investigation, given its capability to leverage co-evolution signals encoded in aligned homologous sequences. Moreover, the optimized MSA Transformer outperforms existing methods and exhibits enhanced generalization power, leading to a notable improvement in predicting changes in protein stability resulting from point mutations. Code and data are available at https://github.com/RitAreaSciencePark/PLM4Muts.

bioinformatics↗

Protein family annotation for the Unified Human Gastrointestinal Proteome by DPCfam clustering

Technological advances in massively parallel sequencing have led to an exponential growth in the number of known protein sequences. Much of this growth originates from metagenomic projects producing new sequences from environmental and clinical samples. The Unified Human Gastrointestinal Proteome (UHGP) catalogue is one of the most relevant metagenomic datasets with applications ranging from medicine to biology. However, the lack of sequence annotation impairs its usability. This work aims to produce a family classification of UHGP sequences to facilitate downstream structural and functional annotation. This is achieved through the release of the DPCfam-UHGP50 dataset containing 10,778 putative protein families generated using DPCfam clustering, an unsupervised pipeline grouping sequences into multi-domain architectures. DPCfam-UHGP50 considerably improves family coverage at protein and residue levels compared to the manually curated repository Pfam. It is our hope that DPCfam-UHGP50 will foster future discoveries in the field of metagenomics of the human gut by the release of a FAIR-compliant database easily accessible via a searchable web server and Zenodo repository.

bioinformatics↗

The geometry of hidden representations of protein language models

Protein language models (pLMs) transform their input into a sequence of hidden representations whose geometric behavior changes across layers. Looking at fundamental geometric properties such as the intrinsic dimension and the neighbor composition of these representations, we observe that these changes highlight a pattern characterized by three distinct phases. This phenomenon emerges across many models trained on diverse datasets, thus revealing a general computational strategy learned by pLMs to reconstruct missing parts of the data. These analyses show the existence of low-dimensional maps that encode evolutionary and biological properties such as remote homology and structural information. Our geometric approach sets the foundations for future systematic attempts to understand the space of protein sequences with representation learning techniques.

bioinformatics↗