Search bioRxiv⌕ Search

Biology subjects

Alvarez Salmoral, D.

Publications and source records attributed to Alvarez Salmoral, D..

4 recordsLinked to original sources

Integrating computational protein structure predictions and genetic dependencies yields an atlas of human multi-protein complexes (AHMPC)

Knowledge of which proteins interact to form functional complexes in cells is essential for understanding molecular mechanisms in biology. Structure prediction methods recently allowed to compute the Human Interactome of likely binary protein interactions. We combine computational predictions with orthogonal functional data from the Dependency Map that estimate the correlation between genetic dependencies and vulnerabilities of the corresponding gene pairs. This revealed groups of proteins that likely form larger complexes. Clustering analysis followed by AlphaFold3 multi-protein complex predictions and AlphaBridge analysis provided the basis to construct and atlas of human multi-protein complexes (AHMPC), currently encompassing 354 high-confidence predicted multi-protein complexes. These include well-known assemblies and new ones - such as a complex involving SYS1, JTB, and ARFRP1 that we validate experimentally, suggesting an unexpected role of JTB in Golgi traficking. To enable the research community to explore the AHMPC and enable further discovery, we cluster all complexes using functional and disease-related embeddings, demonstrate how structured prompts allow validation by large language models (LLMs), and make all structures and analysis available online as an open resource at https://ahmpc.eu/.

biochemistry↗

DSSP 4: FAIR annotation of protein secondary structure

Protein secondary structure annotation is essential for understanding protein architecture, serving as a cornerstone for structural classification, alignment, visualisation, and machine learning applications. The Define Secondary Structure of Proteins (DSSP) algorithm has long been the standard for assigning secondary structure elements such as -helices, {beta}-sheets, and loops in protein models. Here, we introduce DSSP version 4, which recapitulates DSSP functionality in a modern computational framework, extending also to the detection of left-handed {kappa}-helices (Poly-Proline II helices). To align with the FAIR principles (Findable, Accessible, Interoperable, Reusable), DSSP 4 adopts mmCIF as its primary input and output format, while retaining compatibility with legacy PDB and DSSP formats. We applied this updated tool to analyse the distribution of secondary structure elements across the Protein Data Bank (PDB) differentiating structures from diverse experimental methods, revealing insights into the prevalence and length of secondary structure elements, including the newly annotated {kappa}-helices. The DSSP 4 software, databank, and server are freely accessible from https://pdb-redo.eu/dssp, ensuring broad utility and interoperability in structural biology research.

bioinformatics↗

Disentangling the CHAOS of intrinsic disorder in human proteins

Most proteins consist of both folded domains and Intrinsically Disordered Regions (IDRs). However, the widespread occurrence of intrinsic disorder in human proteins, along with its characteristics, is often overlooked by the broader communities of structural and molecular biologists. Building on the MobiDB database of intrinsic disorder in proteins, here we develop a comprehensive dataset (Comprehensive analysis of Human proteins And their disOrdered Segments - CHAOS). We implement internally consistent definitions of disordered regions, and annotate general characteristics such as cellular location, essentiality, post-translational modifications, and predicted pathogenicity. Further, we cross-reference to structure predictions from AlphaFold. We find that most human proteins contain at least one disordered region, predominantly located at the protein termini. IDRs are less hydrophobic, enriched in post-translational modifications, and mutations in IDRs are predicted to be less pathogenic than in non-IDRs. Additionally, we discovered that proteins residing in different cellular locations possess distinct disorder profiles. Finally, the predicted AlphaFold models of proteins in CHAOS suggest that disordered regions and proteins are often predicted to adopt secondary structure. Hereby we enhance the visibility and understanding of intrinsic disorder in human proteins. Key messagesO_LIFour out of five human proteins contain one or more intrinsically disordered regions (IDRs). C_LIO_LIHalf of the IDRs are located at protein termini, but three quarters of all human proteins contain a terminal IDR. C_LIO_LIThe amount and location of disordered regions differs throughout cellular compartments. C_LIO_LIOne in five missense mutations in IDRs are likely pathogenic. C_LIO_LIAlphaFold predicts secondary structure elements within intrinsically disordered regions and fully disordered proteins. C_LI

bioinformatics↗

AlphaBridge: tools for the analysis of predicted macromolecular complexes

Artificial intelligence (AI)-powered protein structure prediction has transformed how scientists explore macromolecular function. AI-based prediction of macromolecular complexes is increasingly used for evaluating the likelihood of proteins forming complexes with other proteins, nucleic acids, lipids, sugars, or small-molecule ligands. Efficient tools are needed to evaluate these predicted models. We introduce an approach based on combining the confidence metrics of AlphaFold3 to enable clustering of sequence motifs participating in binary interactions and subsequently in 3D interfaces of complexes. Interaction interfaces within confidence limits are finally visualised in 2D using chord diagrams and network graphs. The analysis and visualisation are implemented in a web tool, which links them with interactive graphics and summary tables of predicted interfaces and intermolecular interactions, including confidence scores. Finally, we demonstrate real-life examples of how AlphaBridge is used for providing an efficient way to assess and validate predicted protein complexes and interfaces. The reproducible, objective and automated procedures we present provide a straightforward critical assessment of structure prediction of biomolecular complexes, that should be consulted before conducting more resource-intensive analyses.

bioinformatics↗