Search bioRxiv⌕ Search

Biology subjects

Sellner, M. S.

Publications and source records attributed to Sellner, M. S..

3 recordsLinked to original sources

PanScreen: A Comprehensive Approach to Off-Target Liability Assessment

Drug development projects are getting increasingly more expensive while their success rate is stagnating. Safety issues attributed to off-target binding represent a major reason for the failure of new drugs. Besides desired on-target binding, small molecules may interact with off-targets, triggering adverse effects. Therefore, the development of novel methods for early recognition of such issues that are resource-efficient and cost-effective becomes vital. Here, we introduce PanScreen, an online platform for the automated assessment of off-target liabilities. PanScreen combines structure-based modeling techniques with state-of-the-art deep learning methods to not only predict accurate binding affinities but also give insight into potential modes of action. We show that the predictions are approaching experimental accuracy found in public datasets and that the same technology can also be used for other research areas, such as drug repurposing. Such fast and inexpensive methods allow researchers to test not only drug candidates, but all small molecules that might come into contact with a human organism for potential safety concerns very early in the development process. PanScreen is publicly available at www.panscreen.ch.

bioinformatics↗

Enhancing Ligand-Based Virtual Screening with 3D Shape Similarity via a Distance-Aware Transformer Model

Following the assumption that chemically similar molecules exhibit similar biologcial properties, ligand-based virtual screening can be a valuable starting point in drug discovery projects. While 2D-based similarity metrics generally focus on similar scaffolds or substructures, 3D-based methods can capture the shape of a molecule, allowing for the identification of compounds with different scaffolds. We recently published a proof-of-concept study which demonstrated how a Transformer model can be adapted to preserve 2D similarities in latent space in the form of Euclidean distances. In this work, we extend this research and prove that the approach can be adapted to 3D similarities. We use pharmacophore-based shape similarity as 3D similarity measure. We show that the model is able to enrich the predicted most similar hits with compounds with different scaffolds that are indeed similar in 3D space. Whereas classical pharmacophore- or shape-based 3D similarity methods rely on expensive alignment processes, in our approach, we identify similar compounds directly by the Euclidean distances in latent space. This enables for the first time the 3D screening of ultra-large databases with high efficiency.

bioinformatics↗

Quality Matters: Deep Learning-Based Analysis of Protein-Ligand Interactions with Focus on Avoiding Bias

The efficient and accurate prediction of protein-ligand binding affinities is an extremely appealing yet still unresolved goal in computational pharmacy. In recent years, many scientists have taken advantage of the remarkable progress of deep learning and applied it to address this issue. Despite all the advances in this field, there is increasing evidence that the typically applied validation of these methods is not suitable for medicinal chemistry applications. This work assesses the importance of dataset quality and proper dataset splitting techniques demonstrated on the example of the PDBbind dataset. We also introduce a new tool for the analysis of protein-ligand complexes, called po-sco. Po-sco allows the extraction of interaction information with much higher detail and comprehensibility than the tools available to date. We trained a transformer-based deep learning model to generate protein-ligand interaction fingerprints that can be utilized for downstream predictions, such as binding affinity. When using po-sco, this model generated predictions that were superior to those based on commonly used PLIP and ProLIF tools. We also demonstrate that the quality of the dataset is more important than the number of data points and that suboptimal dataset splitting can lead to a significant overestimation of model performance.

bioinformatics↗