Search bioRxiv⌕ Search

Biology subjects

Huetter, J.-C.

Publications and source records attributed to Huetter, J.-C..

5 recordsLinked to original sources

Spatial Decoding of Tertiary Lymphoid Structure Maturation in Non-Small Cell Lung Cancer Using Deep Neural Networks

Understanding the role of tertiary lymphoid structures (TLS) is crucial in non-small cell lung cancer (NSCLC), as they are associated with patient prognosis and treatment outcomes. Specific cellular ecosystems that originate anti-tumor activity or predict immunotherapy response remain poorly characterized. To this end, we developed a high-resolution, multimodal spatial atlas jointly profiling transcriptomics, proteomics, and histology to characterize TLS maturation in NSCLC alongside secondary lymph organs as a baseline. Using this atlas, we proposed a pathologist-in-the-loop framework that combines a variational graph autoencoder (VGAE) with diffusion pseudotime to refine human expert annotations and characterize TLS maturation. These spatial molecular representations were extended to H&E whole-slide images via a vision transformer-based foundation model. Next, we resolved cellular composition, spatial organization, and cell-cell interactions within these data and defined two divergent spatial ecosystems. Clinical evidence suggests that these ecosystems are associated with distinct patient outcomes: a mature germinal center niche with favorable prognoses, and a tumor-macrophage-fibroblast niche with unfavorable prognoses. In summary, our work decodes key components of TLS heterogeneity, identifies hallmark spatial patterns involved in NSCLC adaptive immunity, and provides a framework for translating spatial omics insights into clinical applications.

genomics↗

Inferring cellular communication through mapping cells in space using Tangram 2

Cell-to-cell communication (CCC) shapes development, immunity, and disease, yet current spatial transcriptomics (SRT) platforms rarely achieve both single-cell resolution and high-quality transcriptome-wide coverage to accurately characterize CCC in the tissue microenvironment. Tangram2 bridges this gap by integrating single-cell RNA sequencing (scRNA-seq) with SRT to identify genes whose expression changes as a function of neighboring cell types. By accurately mapping cells in tissue space, Tangram2 disentangles interaction-driven transcriptional shifts from intrinsic identity markers, yielding mechanistic CCC maps. Validation across diverse settings--including Slide-tags, co-cultured experiments, human lymph nodes, and simulations--demonstrates high accuracy. Applied to triple-negative breast cancer (TNBC) and cutaneous squamous cell carcinoma (cSCC), Tangram2 recapitulates known biology and uncovers new hypotheses, including various immunosuppressive mechanisms in TNBC and a macrophage-regulatory T-cell circuit associated with survival in cSCC.

genomics↗

A high-throughput phenotypic screen combined with an ultra-large-scale deep learning-based virtual screening reveals novel scaffolds of antibacterial compounds

The proliferation of multi-drug-resistant bacteria underscores an urgent need for novel antibiotics. Traditional discovery methods face challenges due to limited chemical diversity, high costs, and difficulties in identifying structurally novel compounds. Here, we explore the integration of small molecule high-throughput screening with a deep learning-based virtual screening approach to uncover new antibacterial compounds. Leveraging a diverse library of nearly 2 million small molecules, we conducted comprehensive phenotypic screening against a sensitized Escherichia coli strain that, at a low hit rate, yielded thousands of hits. We trained a deep learning model, GNEprop, to predict antibacterial activity, ensuring robustness through out-of-distribution generalization techniques. Virtual screening of over 1.4 billion compounds identified potential candidates, of which 82 exhibited antibacterial activity, illustrating a 90X improved hit rate over the high-throughput screening experiment GNEprop was trained on. Importantly, a significant portion of these newly identified compounds exhibited high dissimilarity to known antibiotics, indicating promising avenues for further exploration in antibiotic discovery.

bioinformatics↗

Multi-ContrastiveVAE disentangles perturbation effects in single cell images from optical pooled screens

Optical pooled screens (OPS) enable comprehensive and cost-effective interrogation of gene function by measuring microscopy images of millions of cells across thousands of perturbations. However, the analysis of OPS data still mainly relies on hand-crafted features, even though these are difficult to deploy across complex data sets. This is because most unsupervised feature extraction methods based on neural networks (such as auto-encoders) have difficulty isolating the effect of perturbations from the natural variations across cells and experimental batches. Here, we propose a contrastive analysis framework that can more effectively disentangle the phenotypes caused by perturbation from natural cell-cell heterogeneity present in an unperturbed cell population. We demonstrate this approach by analyzing a large data set of over 30 million cells imaged across more than 5, 000 genetic perturbations, showing that our method significantly outperforms traditional approaches in generating biologically-informative embeddings and mitigating technical artifacts. Furthermore, the interpretable part of our model distinguishes perturbations that generate novel phenotypes from the ones that only shift the distribution of existing phenotypes. Our approach can be readily applied to other small-molecule and genetic perturbation data sets with highly multiplexed images, enhancing the efficiency and precision in identifying and interpreting perturbation-specific phenotypic patterns, paving the way for deeper insights and discoveries in OPS analysis.

bioinformatics↗

Biological Cartography: Building and Benchmarking Representations of Life

1The continued scaling of genetic perturbation technologies combined with high-dimensional assays such as cellular microscopy and RNA-sequencing has enabled genome-scale reverse-genetics experiments that go beyond single-endpoint measurements of growth or lethality. Datasets emerging from these experiments can be combined to construct perturbative "maps of biology", in which readouts from various manipulations (e.g., CRISPR-Cas9 knockout, CRISPRi knockdown, compound treatment) are placed in unified, relatable embedding spaces allowing for the generation of genome-scale sets of pairwise comparisons. These maps of biology capture known biological relationships and uncover new associations which can be used for downstream discovery tasks. Construction of these maps involves many technical choices in both experimental and computational protocols, motivating the design of benchmark procedures to evaluate map quality in a systematic, unbiased manner. Here, we (1) establish a standardized terminology for the steps involved in perturbative map building, (2) introduce key classes of benchmarks to assess the quality of such maps, (3) construct maps from four genome-scale datasets employing different cell types, perturbation technologies, and data readout modalities, (4) generate benchmark metrics for the constructed maps and investigate the reasons for performance variations, and (5) demonstrate utility of these maps to discover new biology by suggesting roles for two largely uncharacterized genes. 2 Author SummaryWith the proliferation of genetic perturbation, laboratory robotics, computer vision and sequencing technologies, a growing number of researchers are producing datasets that capture digital readouts of cellular responses to genetic perturbations at the full-genome-scale. Since each of these efforts utilizes different cellular models, experimental approaches, terminology, code bases, analysis methods and quality metrics, it is exceptionally difficult to reason through the pros and cons of possible design choices or even discuss the primary considerations when embarking on such an endeavor. These datasets can be powerful discovery tools to look at known biological relationships and uncover new associations in an unbiased manner, but only when paired with a computational pipeline to assemble the data into a digestible format. Moreover, there is great promise in looking across these data to highlight commonalities and differences that may be attributed to experimental or analytical approaches or the biological context. Therefore, a unified framework is necessary to align this nascent field and speed progress in assessing technologies and methods. In this work we define a unified framework for building and benchmarking these perturbative maps, benchmark four different datasets assembled into 18 different maps, explore the impact of different design decisions and demonstrate how these maps can be used to elucidate gene functions. The framework we propose highlights the necessary steps for building any such map - embedding, filtering, aligning, aggregating and relating the data across perturbations. For benchmarking, we propose two main types of metrics and give examples which highlight the impact of different processing pipelines. Finally, we explore these maps to demonstrate their utility for confirming known biological relationships and nominating annotations for genes with unknown function. We expect that this work will positively impact the nascent field of perturbative map building by enabling easier comparisons within and between technologies and methods through a shared language. Additionally, the associated code base is openly available and flexible enough to be easily extended with new methods, so we hope that it will become a resource for future researchers working on developing both laboratory and computational methodology. While there are too many confounding variables to make recommendations on the strengths of different technologies and cellular models at this time, highlighting that fact may prompt studies designed with the goal of directly comparing methods while holding other confounding variables fixed. Moreover, as the number of perturbative maps grows, the field will naturally consider the advantages of combining maps across modalities and the framework provided here can also help guide the evaluation of those efforts.

bioinformatics↗