Search bioRxiv⌕ Search

Biology subjects

Woillard, J.-B.

Publications and source records attributed to Woillard, J.-B..

2 recordsLinked to original sources

A unified benchmark of synthetic data generation for clinical transcriptomic cancer cohorts

Achieving a trade-off between biological utility and patient privacy remains a key challenge for secure data sharing when applying transcriptomic clinical datasets to artificial intelligence in precision oncology. Here, we introduce the first benchmarking study tailored to high-dimensional clinical transcriptomic cancer data, comparing synthetic data generation methods across three clinical cancer trials. Our framework, SynOmicsBench, combines standardized preprocessing with multidimensional evaluation, prioritizing downstream biological validation alongside statistical fidelity and attack-based privacy assessment. Results indicate that no single method dominated all dimensions, with Gaussian Copula achieving the most balanced performance, followed by Avatar, demonstrating that metric-based similarity alone is insufficient to ensure preservation of higher-order molecular dependencies. Synthetic data consistently reproduced biomedical signal directionality but with attenuated effect sizes and inter-replicate variability, supporting hypothesis generation when multi-seed synthesis is adopted. Collectively, this framework provides a reproducible decision-support tool for method selection and promotes biologically informed, privacy-aware adoption of synthetic data in precision oncology.

bioinformatics↗

Model Ensembling and Machine Learning Approaches to Predict the First Dose of Amoxicillin in Intensive Care

A priori model informed precision dosing (MIPD) recommends an appropriate first dose based solely on the patients covariates enabling faster target attainment without required concentration measurements. Population pharmacokinetic model ensembling and machine learning (ML) approaches were developed and evaluated to predict a first dose of amoxicillin in intensive care. Following a bibliographic review, a virtual patient population was simulated based on cohorts from four published adult amoxicillin PopPK models. Model-development cohorts were reproduced, and steady-state trough concentrations were simulated using cohort-specific dosing regimens. As reference methods, weighted model ensembling (WME) and classification tree (CT)-informed ensembling were implemented. Two novel ensembling strategies were developed: regression tree (RT)-informed ensembling, using RT to predict the log individual prediction/observation ratio, and factor analysis of mixed data (FAMD), assigning model weights based on patient similarity to original model cohorts. In parallel, four ML algorithms (support vector machine, k-nearest neighbors, random forest, and XGBoost) were trained to predict the dose achieving target concentrations based on covariates and dosing scheme. All approaches were compared with single-model PopPK dosing, standard dosing, and a nomogram, and externally validated using clinical data. Most MIPD methods outperformed standard dosing. On simulated data, ensembling (30-42 % correct predictions) and ML (36-39 %) exceeded single-model approaches (14-32 %). RT-informed and FAMD ensembling improved performance by 6-10 % over uninformed ensembling on clinical data. In clinical patients receiving continuous infusion, ensembling further improved performance, with FAMD achieving 49 % correct predictions. ML-based ensembling eliminates the need for model selection and increase target attainment.

pharmacology and toxicology↗