Search bioRxiv⌕ Search

Biology subjects

Herrera, L. J.

Publications and source records attributed to Herrera, L. J..

2 recordsLinked to original sources

The SEA-AD DREAM Challenge: Community benchmarking human and AI agent solutions for Alzheimer's disease neuropathology prediction from single-nucleus transcriptomics

Single-nucleus transcriptomic atlases offer an unprecedented opportunity to connect cellular molecular states with Alzheimer's disease (AD) neuropathology, but whether these profiles encode reproducible, predictive information about pathological burden remains unclear. We present the SEA-AD DREAM Challenge, an open, international, model-to-data competition built on the Seattle Alzheimer's Disease Brain Cell Atlas to predict Alzheimer's disease neuropathological severity from single-nucleus RNA-sequencing data. Participants developed containerized models to predict categorical neuropathological staging, including overall Alzheimer's disease neuropathologic change, Braak stage, Thal phase, and CERAD score, as well as quantitative amyloid-{beta} and phospho-tau burden measured by 6E10 and AT8 immunohistochemistry. Across 17 eligible teams from 15 countries, the crowdsourcing framework enabled systematic comparison of diverse computational approaches and surfaced a broad landscape of modeling strategies and candidate predictive features. Top-performing methods achieved near-perfect prediction of categorical staging, with the best submission reaching a quadratic weighted kappa of 1.0 for the Overall AD Neuropathological Change score (ADNC), and competitive prediction of quantitative pathological burden in held-out data, with a best concordance correlation coefficient of 0.48. Post hoc perturbation analyses revealed that top categorical-stage predictions relied heavily on donor-level metadata-driven signals rather than transcriptomic features, whereas quantitative pathology prediction was more robust and supported by transcriptomic and cell-type-associated features with potential biological relevance to AD progression. The challenge also introduced the first AI Agent Track in a DREAM Challenge, providing an early benchmark for autonomous and human-guided agentic model development in single-cell neuroscience. This work demonstrates that single-nucleus transcriptomes encode substantial information about Alzheimer's disease pathology, establishes a reproducible benchmark for molecular neuropathology prediction, and highlights critical principles for designing privacy-preserving, leakage-aware community challenges using deeply phenotyped human brain data.

neuroscience↗

Synthetic whole-slide image tile generation with gene expression profiles infused deep generative models

The acquisition of multi-modal biological data for the same sample, such as RNA sequencing and whole slide imaging (WSI), has increased in recent years, enabling studying human biology from multiple angles. However, despite these emerging multi-modal efforts, for the majority of studies only one modality is typically available, mostly due to financial or logistical constraints. Given these difficulties, multi-modal data imputation and multi-modal synthetic data generation are appealing as a solution for the multi-modal data scarcity problem. Currently, most studies focus on generating a single modality (e.g. WSI), without leveraging the information provided by additional data modalities (e.g. gene expression profiles). In this work, we propose an approach to generate WSI tiles by using deep generative models infused with matched gene expression profiles. First, we train a variational autoencoder (VAE) that learns a latent, lower dimensional representation of multi-tissue gene expression profiles. Then, we use this representation to infuse generative adversarial networks (GAN) that generate lung and brain cortex tissue tiles, resulting in a new model that we call RNA-GAN. Tiles generated by RNA-GAN were preferred by expert pathologists in comparison to tiles generated using traditional GANs and in addition, RNA-GAN needs fewer training epochs to generate high-quality tiles. Finally, RNA-GAN was able to generalize to gene expression profiles outside of the training set, showing imputation capabilities. A web-based quiz is available for users to play a game distinguishing real and synthetic tiles: https://rna-gan.stanford.edu/ and the code for RNA-GAN is available here: https://github.com/gevaertlab/RNA-GAN.

bioinformatics↗