Search bioRxiv⌕ Search

Biology subjects

Medina-Ortiz, D.

Publications and source records attributed to Medina-Ortiz, D..

4 recordsLinked to original sources

Protein language models accelerate the discovery of Plastic-Degrading Enzymes

Plastic pollution presents a critical environmental challenge, necessitating innovative and sustainable solutions. In this context, biodegradation using microorganisms and enzymes offers an environmentally friendly alternative. This work introduces an AI-driven frame-work that integrates machine learning (ML) and generative models to accelerate the discovery and design of plastic-degrading enzymes. By leveraging pre-trained protein language models and curated datasets, we developed seven ML-based binary classification models to identify enzymes targeting specific plastic substrates, achieving an average accuracy of 89%. The framework was applied to over 6,000 enzyme sequences from the RemeDB to classify enzymes targeting diverse plastics, including PET, PLA, and Nylon. Besides, generative learning strategies combined with trained classification models in this work were applied for de novo generation of PET-degrading enzymes. Structural bioinformatics validated potential candidates through in-silico analysis, highlighting differences in physicochemical properties between generated and experimentally validated enzymes. Moreover, generated sequences exhibited lower molecular weights and higher aliphatic indices, features that may enhance interactions with hydrophobic plastic substrates. These findings highlight the utility of AI-based approaches in enzyme discovery, providing a scalable and efficient tool for addressing plastic pollution. Future work will focus on experimental validation of promising candidates and further refinement of generative strategies to optimize enzymatic performance.

bioengineering↗

Peptipedia v2.0: A peptide sequence database and user-friendly web platform. A major update

In recent years, peptides have gained significant relevance due to their therapeutic properties. The surge in peptide production and synthesis has generated vast amounts of data, enabling the creation of comprehensive databases and information repositories. Advances in sequencing techniques and artificial intelligence have further accelerated the design of tailor-made peptides. However, leveraging these techniques requires versatile and continuously updated storage systems, along with tools that facilitate peptide research and the implementation of machine learning for predictive systems. This work introduces Peptipedia v2.0, one of the most comprehensive public repositories of peptides, supporting biotechnological research by simplifying peptide study and annotation. Peptipedia v2.0 has expanded its collection by over 45% with peptide sequences that have reported biological activities. The functional biological activity tree has been revised and enhanced, incorporating new categories such as cosmetic and dermatological activities, molecular binding, and anti-ageing properties. Utilizing protein language models and machine learning, more than 90 binary classification models have been trained, validated, and incorporated into Peptipedia v2.0. These models exhibit average sensitivities and specificities of 0.877 {+/-} 0.0530 and 0.873 {+/-}0.054, respectively, facilitating the annotation of more than 3.6 million peptide sequences with unknown biological activities, also registered in Peptipedia v2.0. Additionally, Peptipedia v2.0 introduces description tools based on structural and ontological properties and user-friendly machinelearning tools to facilitate the application of machine-learning strategies to study peptide sequences. Peptipedia v2.0 is accessible under the Creative Commons CC BY-NC-ND 4.0 license at https://peptipedia.cl/.

bioinformatics↗

Interpretable and explainable predictive machine learning models for data-driven protein engineering

Protein engineering using directed evolution and (semi)rational design has emerged as a powerful strategy for optimizing and enhancing enzymes or proteins with desired properties. Integrating artificial intelligence methods has further enhanced and accelerated protein engineering through predictive models developed in data-driven strategies. However, the lack of explainability and interpretability in these models poses challenges. Explainable Artificial Intelligence addresses the interpretability and explainability of machine learning models, providing transparency and insights into predictive processes. Nonetheless, there is a growing need to incorporate explainable techniques in predicting protein properties in machine learning-assisted protein engineering. This work explores incorporating explainable artificial intelligence in predicting protein properties, emphasizing its role in trustworthiness and interpretability. It assesses different machine learning approaches, introduces diverse explainable methodologies, and proposes strategies for seamless integration, improving trust-worthiness. Practical cases demonstrate the explainable models effectiveness in identifying DNA binding proteins and optimizing Green Fluorescent Protein brightness. The study highlights the utility of explainable artificial intelligence in advancing computationally assisted protein design, fostering confidence in model reliability.

bioinformatics↗

RUDEUS, a machine learning classification system to study DNA-Binding proteins

DNA-binding proteins are essential in different biological processes, including DNA replication, transcription, packaging, and chromatin remodelling. Exploring their characteristics and functions has become relevant in diverse scientific domains. Computational biology and bioinformatics have assisted in studying DNA-binding proteins, complementing traditional molecular biology methods. While recent advances in machine learning have enabled the integration of predictive systems with bioinformatic approaches, there still needs to be generalizable pipelines for identifying unknown proteins as DNA-binding and assessing the specific type of DNA strand they recognize. In this work, we introduce RUDEUS, a Python library featuring hierarchical classification models designed to identify DNA-binding proteins and assess the specific interaction type, whether single-stranded or double-stranded. RUDEUS has a versatile pipeline capable of training predictive models, synergizing protein language models with supervised learning algorithms, and integrating Bayesian optimization strategies. The trained models have high performance, achieving a precision rate of 95% for DNA-binding identification and 89% for discerning between single-stranded and doublestranded interactions. RUDEUS includes an exploration tool for evaluating unknown protein sequences, annotating them as DNA-binding, and determining the type of DNA strand they recognize. Moreover, a structural bioinformatic pipeline has been integrated into RUDEUS for validating the identified DNA strand through DNA-protein molecular docking. These comprehensive strategies and straightforward implementation demonstrate comparable performance to high-end models and enhance usability for integration into protein engineering pipelines.

bioinformatics↗