Search bioRxiv⌕ Search

Biology subjects

Mav, D.

Publications and source records attributed to Mav, D..

2 recordsLinked to original sources

Leveraging Targeted Gene Sets and Neural Networks for Zebrafish Transcriptome Extrapolation in High-Throughput Toxicogenomics

BackgroundZebrafish (Danio rerio) are a powerful vertebrate model for developmental toxicology and chemical safety assessment, yet large-scale transcriptomics in zebrafish remains limited by cost and data heterogeneity. Targeted transcriptomics offers a cost-effective alternative, but gene extrapolation methods tailored to zebrafish have not been systematically developed or evaluated. ObjectivesWhile the S1500+ platform is widely used for toxicogenomics research with rat, mouse, and human cell lines as model systems, its use in zebrafish has been limited due to data scarcity and lack of suitable bioinformatics approaches for analysis of such data. To that end, we sought to (i) curate a large zebrafish transcriptomic training data resource, and (ii) evaluate multiple machine learning strategies for reconstructing unmeasured transcriptome-wide expression profiles for data originating from the zebrafish-specific reduced representation gene set ("Zf S1500+"). MethodsWe assembled 14,924 zebrafish RNA-Seq samples covering 21,930 genes across 1,246 studies. Using the Zf S1500+ gene subset (3,062 genes), we trained and tested three extrapolation approaches: principal components regression (PCR), a locally weighted extension of PCR (PCR+), and a neural network mixture-of-experts model (NN-MoE). Model performance was assessed using mean absolute error (MAE), mean squared regression error (MSRE), and weighted variants of these metrics. ResultsExtrapolation performance using the baseline approach was strongly influenced by tissue and developmental context, with within-tissue models outperforming cross-tissue models. Errors were lowest when training and testing were conducted within the same tissue or between developmentally related tissues. Both PCR+ and NN-MoE improved upon the baseline PCR approach, with NN-MoE reducing average MAE by [~]20% and MSRE by [~]17%. Importantly, extrapolation remained reliable for the majority of genes, even when limiting output to high-confidence predictions using an empirical MAE threshold. ConclusionsWe demonstrate that targeted transcriptomics can be effectively extended to zebrafish, enabling robust transcriptome-wide extrapolation at reduced cost. The NN-MoE method provided the most substantial gains, highlighting the value of non-linear and ensemble modeling in heterogeneous datasets. These results establish a scalable framework for zebrafish toxicogenomics and suggest that accuracy will continue to improve with larger, better-annotated datasets, paving the way for broader application in chemical safety assessments.

bioinformatics↗

Characterization and optimization of variability in a human colonic epithelium culture model

Animal models have historically been poor preclinical predictors of gastrointestinal (GI) directed therapeutic efficacy and drug-induced GI toxicity. Human stem and primary cell-derived culture systems are a major focus of efforts to create biologically relevant models that enhance preclinical predictive value of intestinal efficacy and toxicity. The inherent variability in stem-cell-based complex cultures makes development of useful models a challenge; the stochastic nature of stem-cell differentiation interferes with the ability to build and validate robust, reproducible assays that query drug responses and pharmacokinetics. In this study, we aimed to characterize and reduce potential sources of variability in a complex stem cell-derived intestinal epithelium model, termed RepliGut(R) Planar, across cells from multiple human donors, cell lots, and passage numbers. Assessment criteria included barrier formation and integrity, gene expression, and cytokine responses. Gene expression and culture metric analyses revealed that controlling for stem/progenitor-cell passage number reduces variability and maximizes physiological relevance of the model. After optimizing passage number, donor-specific differences in cytokine responses were observed in a case study, suggesting biologic variability is observable in cell cultures derived from multiple human sources. Our findings highlight key considerations for designing assays that can be applied to additional primary-cell derived systems, as well as establish utility of the RepliGut(R) Planar platform for robust development of human-predictive drug-response assays.

cell biology↗