Search bioRxiv⌕ Search

Biology subjects

Tilewale, A.

Publications and source records attributed to Tilewale, A..

2 recordsLinked to original sources

Boltz2-Notebook: An Interactive Google Colab Platform for Diffusion-Based Biomolecular Structure and Binding Affinity Prediction using the Boltz2 model.

Recent advances in deep learning-based structure prediction, including AlphaFold3 and the open-source Boltz model family, have extended biomolecular modeling to joint prediction of protein-ligand, protein-nucleic acid, and multi-chain complexes with binding-affinity estimation. Boltz-2 is among the most feature-complete of these open models, but its practical use requires a local CUDA-capable GPU, command-line execution, and manually authored YAML configuration files, limiting accessibility for researchers without dedicated computational infrastructure. We developed Boltz2-Notebook, a Colab-native interface comprising four integrated stages - automated environment setup, interactive parameter-to-YAML generation, execution management, and automated confidence and affinity visualization - together with a manifest-driven batch mode for multi-target screening. All modelling capabilities are inherited unmodified from Boltz-2; Boltz2-Notebook's contributions are limited to accessibility, input construction, and workflow automation. Independent of the software, we curated a benchmark of 317 protein-ligand pairs (122 proteins, 277 ligands) from BindingDB and predicted binding affinity in triplicate using the Boltz-2 command-line engine on high-performance computing infrastructure. Predicted and experimental pIC50 values showed moderate correlation (Pearson r = 0.609 [95% CI 0.540-0.675]; Spearman {rho} = 0.625; R2 = 0.371; MAE = 0.968 pIC50 units), with high triplicate reproducibility (pairwise r = 0.97) but a systematic compression of the predicted affinity range and no measurable relationship between Boltz-2's self-reported confidence metrics and prediction accuracy. Boltz2-Notebook is freely available as open-source software and provides external, reproducible evidence - including a specific confidence-calibration limitation - relevant to interpreting Boltz-2 affinity predictions responsibly.

bioinformatics↗

Toward Digital-Twin-Enabled Bioprocess Monitoring: Fault-Inclusive Soft Sensing of Penicillin Concentration Under Process Deviations

Data-driven soft sensors can estimate fermentation product concentrations from routinely recorded process variables, but strong performance during normal operation does not establish reliability during process deviations. This study evaluated current-time penicillin-concentration estimation using 100 simulated IndPenSim batches comprising 90 normal-operation batches and ten documented deviation batches, with 113,935 observations in total. Initial normal-trained models and five follow-up experiments examined complete-batch validation, dependence on batch-progress features, phase-specific error, model-family comparisons, early out-of-distribution warnings and empirical prediction ranges. These analyses motivated a matched comparison between normal-only and fault-inclusive HistGradientBoosting regressors using 36 current and causal history-based inputs. Normal performance was evaluated in five regime-balanced complete-batch folds, and deviation performance by leave-one-fault-batch-out evaluation; every tested deviation batch was excluded from its own model fit. Batch-balanced sample weights were used, with a factor of three assigned to permitted deviation batches in fault-inclusive fitting. On held-out deviation batches, pooled RMSE decreased from 3.195 to 2.564 g/L, a reduction of 19.73%, while MAE decreased from 2.076 to 1.441 g/L and R2 increased from 0.8580 to 0.9085. Normal-operation RMSE was nearly unchanged at 1.981 and 1.983 g/L. Eight of ten deviation batches improved. The mean paired fault-batch RMSE difference was -0.7063 g/L, with a descriptive 95% batch-bootstrap interval of -1.2844 to -0.2194 g/L. Improvement was largest within the first-to-last recorded fault-reference window, but late-stage errors persisted. Batch 100 remained poorly predicted, with fault-inclusive RMSE of 6.631 g/L and R2 of -2.4837. A fault-risk classifier, Isolation Forest OOD detector and empirical error ranges offered incomplete reliability information. Fault-inclusive training therefore improved this benchmark on average, but did not establish generalization to unseen fault mechanisms, calibrated safety warnings or deployment in physical fermentation.

bioengineering↗