Search bioRxiv⌕ Search

Biology subjects

Tossou, P.

Publications and source records attributed to Tossou, P..

2 recordsLinked to original sources

Nesso-1: Accelerating Open-Source Binding Affinity Predictions

In this technical report, we introduce NO_SCPLOWESSOC_SCPLOW-1, a coarse-grained cofolding framework for binding- affinity prediction. NO_SCPLOWESSOC_SCPLOW-1 requires[~] 1 second per prediction on a single GPU. This offers more than one order of magnitude speed-up over the leading open-source baseline, Boltz-2, which significantly expands the regions of chemical space that can be explored during high-throughput virtual screening. Importantly, NO_SCPLOWESSOC_SCPLOW-1 matches or surpasses the accuracy of Boltz-2 over the same benchmarks adopted in their study--which we show reflect in-distribution scenarios--as well as over more challenging out-of-distribution data encompassing the OpenBind affinity benchmark and 25 internal biochemical assays. Notably, NO_SCPLOWESSOC_SCPLOW-1 maintains robust predictive accuracy even on assays with extremely low similarity to the training data. Moreover, we highlight examples where NO_SCPLOWESSOC_SCPLOW-1 demonstrates meaningful selectivity, separating the binding affinities of identical compounds between on-targets and related off-targets. Nonetheless, zero-shot generalization to real- world medicinal chemistry remains an inherently challenging task; consequently, we acknowledge specific assays where the models performance is limited. We open-source NO_SCPLOWESSOC_SCPLOW-1: code and weights are available at https://github.com/recursionpharma/nesso

molecular biology↗

MOT: a Multi-Omics Transformer for multiclass classification tumour types predictions

MotivationBreakthroughs in high-throughput technologies and machine learning methods have enabled the shift towards multi-omics modelling as the preferred means to understand the mechanisms underlying biological processes. Machine learning enables and improves complex disease prognosis in clinical settings. However, most multi-omic studies primarily use transcriptomics and epigenomics due to their over-representation in databases and their early technical maturity compared to others omics. For complex phenotypes and mechanisms, not leveraging all the omics despite their varying degree of availability can lead to a failure to understand the underlying biological mechanisms and leads to less robust classifications and predictions. ResultsWe proposed MOT (Multi-Omic Transformer), a deep learning based model using the transformer architecture, that discriminates complex phenotypes (herein cancer types) based on five omics data types: transcriptomics (mRNA and miRNA), epigenomics (DNA methylation), copy number variations (CNVs), and proteomics. This model achieves an F1-score of 98.37% among 33 tumour types on a test set without missing omics views and an F1-score of 96.74% on a test set with missing omics views. It also identifies the required omic type for the best prediction for each phenotype and therefore could guide clinical decisionmaking when acquiring data to confirm a diagnostic. The newly introduced model can integrate and analyze five or more omics data types even with missing omics views and can also identify the essential omics data for the tumour multiclass classification tasks. It confirms the importance of each omic view. Combined, omics views allow a better differentiation rate between most cancer diseases. Our study emphasized the importance of multi-omic data to obtain a better multiclass cancer classification. Availability and implementationMOT source code is available at https://github.com/dizam92/multiomic_predictions.

bioinformatics↗