Search bioRxiv⌕ Search

Biology subjects

De Neve, W.

Publications and source records attributed to De Neve, W..

5 recordsLinked to original sources

Towards Interpretable Multitask Learning for Splice Site and Translation Initiation Site Prediction

In this study, we investigate the effectiveness of multi-task learning (MTL) for handling three bioinformatics tasks: donor splice site prediction, acceptor splice site prediction, and translation initiation site prediction. As the foundation for our MTL approach, we use the SpliceRover model, which has previously been successful in predicting splice sites. While providing benefits such as efficient resource utilization, reduced complexity, and streamlined model management, our findings show that the newly introduced MTL model performs comparably to the SpliceRover model trained separately for each task (single-task models), with a slight decrease in specificity, sensitivity, F1-score, and Matthews Correlation Coefficient (MCC). However, these differences are statistically insignificant (the specificity decreased with 0.0081 for acceptor splice site prediction and the MCC decreased with 0.0264 for TIS prediction), emphasizing the comparable performance of the MTL model. We further analyze the effectiveness of our MTL model using visualization techniques. The outcomes indicate that our MTL model effectively learns the relevant features associated with each task when compared to the single-task models (presence of nucleotides with a higher contribution to donor splice site prediction, polypyrimidine tracts in the upstream of acceptor splice sites, and the Kozak sequence). In conclusion, our results show that the MTL model generalizes well across all three tasks.

bioinformatics↗

Discovering Biomarker Proteins and Peptides for Parkinson's Disease Prognosis Prediction with Machine Learning and Interpretability Methods

Parkinsons disease is a neurodegenerative disorder that affects millions of people worldwide, posing significant challenges for diagnosis and treatment. This study presents a machine learning pipeline for identifying candidate biomarker proteins and peptides from cerebrospinal fluid mass spectrometry (CSF-MS) tests in Parkinsons disease patients. Our pipeline comprises two main stages: (1) model training using mutual information-based feature selection and five different machine learning regressors and (2) identification of candidate biomarkers by combining three types of interpretability methods. Our regression models demonstrated promising effectiveness in predicting the Movement Disorder Society-Unified Parkinsons Disease Rating Scale (MDS-UPDRS) scores, with UPDRS-1 receiving the best predictions, followed by UPDRS-3 and UPDRS-2. Furthermore, our pipeline identified 11 proteins and peptides as potential biomarkers for Parkinsons disease, excluding Levodopa usage which trivially has the most significant impact on the prognosis prediction. Comparisons with four additional pipelines confirmed the effectiveness of our approach in terms of both model performance and biomarker identification. In conclusion, our study presents a comprehensive machine learning pipeline that demonstrates effectiveness in predicting the severity of Parkinsons disease using CSF-MS tests. Our approach also identifies potential biomarkers, which could aid in the development of new diagnostic tools and treatments for patients with Parkinsons disease.

bioinformatics↗

CRISPR-Cas-Docker: Web-based in silico docking and machine learning-based classification of crRNAs with Cas proteins

MotivationCRISPR-Cas-Docker is a web server for in silico docking experiments with CRISPR RNAs (crRNAs) and Cas proteins. This web server aims at providing experimentalists with the optimal crRNA-Cas pair predicted computationally when prokaryotic genomes have multiple CRISPR arrays and Cas systems, as frequently observed in metagenomic data. CRISPR-Cas-Docker provides two methods to predict the optimal Cas protein given a particular crRN sequence: a structure-based method (in silico docking) and a sequence-based method (machine learning classification). For the structure-based method, users can either provide experimentally determined 3D structures of these macromolecules or use an integrated pipeline to generate 3D-predicted structures for in silico docking experiments. ResultsCRISPR-Cas-Docker is an optimized and integrated platform that provides users with 1) 3D-predicted crRNA structures and AlphaFold-predicted Cas protein structures, 2) the top-10 docking models for a particular crRNA-Cas protein pair, and 3) machine learning-based classification of crRNA into its Cas system type. Availability and implementationCRISPR-Cas-Docker is available as an open-source tool under the GNU General Public License v3.0 on GitHub. It is also available as a web server.

bioinformatics↗

In silico optimization of RNA-protein interactions for CRISPR-Cas13-based antimicrobials

RNA-protein interactions are crucial for diverse biological processes. In prokaryotes, RNA-protein interactions enable adaptive immunity through CRISPR-Cas systems. These defense systems utilize CRISPR RNA (crRNA) templates acquired from past infections to destroy foreign genetic elements through crRNA-mediated nuclease activities of Cas proteins. Thanks to the programmability and specificity of CRISPR-Cas systems, CRISPR-based antimicrobials have the potential to be repurposed as new types of antibiotics. Unlike traditional antibiotics, these CRISPR-based antimicrobials can be designed to target specific bacteria and minimize detrimental effects on the human microbiome during antibacterial therapy. Here, we explore the potential of CRISPR-based antimicrobials by optimizing the RNA-protein interactions of crRNAs and Cas13 proteins. CRISPR-Cas13 systems are unique as they degrade specific foreign RNAs using the crRNA template, which leads to non-specific RNase activities and cell cycle arrest. We show that a high proportion of the Cas13 systems have no colocalized CRISPR arrays, and the lack of direct association between crRNAs and Cas proteins may result in suboptimal RNA-protein interactions in the current tools. Here, we investigate the RNA-protein interactions of the Cas13-based systems by curating the validation dataset of Cas13 protein and CRISPR repeat pairs that are experimentally validated to interact, and the candidate dataset of CRISPR repeats that reside on the same genome as the currently known Cas13 proteins. To find optimal CRISPR-Cas13 interactions, we first validate the 3-D structure prediction of crRNAs based on their experimental structures. Next, we test a number of RNA-protein interaction programs to optimize the in silico docking of crRNAs with the Cas13 proteins. From this optimized pipeline, we find a number of candidate crRNAs that have comparable or better in silico docking with the Cas13 proteins of the current tools. This study fully automatizes the in silico optimization of RNA-protein interactions as an efficient preliminary step for designing effective CRISPR-Cas13-based antimicrobials.

bioinformatics↗

Rethinking protein drug design with highly accurate structure prediction of anti-CRISPR proteins

Protein therapeutics play an important role in controlling the functions and activities of disease-causing proteins in modern medicine. Despite protein therapeutics having several advantages over traditional small-molecule therapeutics, further development has been hindered by drug complexity and delivery issues. However, recent progress in deep learning-based protein structure prediction approaches such as AlphaFold opens new opportunities to exploit the complexity of these macro-biomolecules for highly-specialised design to inhibit, regulate or even manipulate specific disease-causing proteins. Anti-CRISPR proteins are small proteins from bacteriophages that counter-defend against the prokaryotic adaptive immunity of CRISPR-Cas systems. They are unique examples of natural protein therapeutics that have been optimized by the host-parasite evolutionary arms race to inhibit a wide variety of host proteins. Here, we show that these Anti-CRISPR proteins display diverse inhibition mechanisms through accurate structural prediction and functional analysis. We find that these phage-derived proteins are extremely distinct in structure, some of which have no homologues in the current protein structure domain. Furthermore, we find a novel family of Anti-CRISPR proteins which are structurally homologous to the recently-discovered mechanism of manipulating host proteins through enzymatic activity, rather than through direct inference. Using highly accurate structure prediction, we present a wide variety of protein-manipulating strategies of anti-CRISPR proteins for future protein drug design.

bioinformatics↗