Search bioRxiv⌕ Search

Biology subjects

Azinfar, L.

Publications and source records attributed to Azinfar, L..

3 recordsLinked to original sources

Tissue-aware elastic net decomposition reveals shared and lineage-specific drug response biomarkers

MotivationComputational models that predict cancer drug response from genomic features are central to biomarker discovery, yet a recent audit found data leakage in 72% of 32 published methods, and complex models offer little interpretability while only modestly exceeding simple baselines under honest evaluation. Tissue lineage is a largely untapped source of legitimate inductive bias, but existing tissue-aware methods neither separate pan-cancer from lineage-specific signal nor report leakage-free performance. ResultsWe introduce the Data Shared Elastic Net (DSEN), a tissue-aware regression that decomposes each drugs model into a shared coefficient block common to all lineages and tissue-specific deviation blocks. Under leakage-free cross-validation across 265 drugs, 1,462 cell lines and 31 tissue lineages, DSEN improved mean squared error over a standard elastic net for 92.5% of drugs (mean 4.95%) while selecting 58% fewer stable shared features. Shared coefficients generalized to held-out tissues (59% tissue-level win rate) and recurrently recovered transferable pathway modules (p53, MAPK), whereas tissue blocks captured lineage markers such as the skin MITF /S100B program. The closest tissue-aware comparator, TG-LASSO, performed worse than the tissue-agnostic baseline (-13.8% mean MSE). Ablation shows tissue-aware modeling helps most when features are scarce, with no single modality dominating. Availability and implementationhttps://github.com/AsiaeeLab/tissue-aware-drug-response. Contactamir.asiaeetaheri@vumc.org

bioinformatics↗

Widespread data leakage inflates performance estimates in cancer drug response prediction

Drug response prediction models guide biomarker discovery, yet their reported accuracy depends on rigorous cross-validation (CV). We show that supervised feature screening applied to all samples before CV, a widespread practice, introduces data leakage that systematically underestimates prediction error. Across 265 drugs and 1,462 cancer cell lines, leakage-free CV raises mean squared error by 16.6% on average, with near-zero feature overlap between leaked and corrected pipelines (mean Jaccard 0.18; 36.2% of drugs share no features). Despite selecting five times more features, the leaked pipeline recovers known drug targets at nearly the same rate as the corrected pipeline, indicating that the inflated feature sets capture statistical artifacts rather than biological signal. A code-level audit of 32 published methods (2017-2024) confirms leakage in 23 (72%), spanning five distinct modes and cited over 3,000 times. The magnitude of inflation from this single leakage mode is comparable to improvement margins typically claimed over elastic net baselines, raising the possibility that some reported advances reflect evaluation artifacts. We provide a leakage taxonomy, audit guide and reference implementation for leakage-free evaluation.

bioinformatics↗

AUC-PR is a More Informative Metric for Assessing the Biological Relevance of In Silico Cellular Perturbation Prediction Models

In silico perturbation models, computational methods which can predict cellular responses to perturbations, present an opportunity to reduce the need for costly and time-intensive in vitro experiments. Many recently proposed models predict high-dimensional cellular responses, such as gene or protein expression to perturbations such as gene knockout or drugs. However, evaluating in silico performance has largely relied on metrics such as R2, which assess overall prediction accuracy but fail to capture biologically significant outcomes like the identification of differentially expressed genes. In this study, we present a novel evaluation framework that introduces the AUC-PR metric to assess the precision and recall of DE gene predictions. By applying this framework to both single-cell and pseudo-bulked datasets, we systematically benchmark simple and advanced computational models. Our results highlight a significant discrepancy between R2 and AUC-PR, with models achieving high R2 values but struggling to identify Differentially expressed genes accurately, as reflected in their low AUC-PR values. This finding underscores the limitations of traditional evaluation metrics and the importance of biologically relevant assessments. Our framework provides a more comprehensive understanding of model capabilities, advancing the application of computational approaches in cellular perturbation research.

bioinformatics↗