Search bioRxiv⌕ Search

Biology subjects

Gander, M.

Publications and source records attributed to Gander, M..

6 recordsLinked to original sources

An integrated transcriptomic cell atlas of human endoderm-derived organoids

Human stem cells can generate complex, multicellular epithelial tissues of endodermal origin in vitro that recapitulate aspects of developing and adult human physiology. These tissues, also called organoids, can be derived from pluripotent stem cells or tissue-resident fetal and adult stem cells. However, it has remained difficult to understand the precision and accuracy of organoid cell states through comparison with primary counterparts, and to comprehensively assess the similarity and differences between organoid protocols. Advances in computational single-cell biology now allow the integration of datasets with high technical variability. Here, we integrate single-cell transcriptomes from 218 samples covering organoids of diverse endoderm-derived tissues including lung, pancreas, intestine, liver, biliary system, stomach, and prostate to establish an initial version of a human endoderm organoid cell atlas (HEOCA). The integration includes nearly one million cells across diverse conditions, data sources and protocols. We align and compare cell types and states between organoid models, and harmonize cell type annotations by mapping the atlas to primary tissue counterparts. To demonstrate utility of the atlas, we focus on intestine and lung, and clarify ontogenic cell states that can be modeled in vitro. We further provide examples of mapping novel data from new organoid protocols to expand the atlas, and showcase how integrating organoid models of disease into the HEOCA identifies altered cell proportions and states between healthy and disease conditions. The atlas makes diverse datasets centrally available, and will be valuable to assess organoid fidelity, characterize perturbed and diseased states, and streamline protocol development.

cell biology↗

CurveCurator: A recalibrated F-statistic to assess, classify, and explore significance of dose-response curves

Dose-response curves are key metrics in pharmacology and biology to assess phenotypic or molecular actions of bioactive compounds in a quantitative fashion. Yet, it is often unclear whether or not a response significantly differs from a curve without regulation, particularly in high-throughput applications or unstable assays. Treating potency and effect size estimates from random and true curves with the same level of confidence can lead to incorrect hypotheses and issues in training machine learning models. Here, we present CurveCurator, an open-source software and interactive dashboard that provides reliable dose-response characteristics by computing p-values and false discovery rates based on a recalibrated F-statistic and a novel thresholding procedure. Application of CurveCurator to large-scale data sets demonstrates its scalable utility across several application areas.

bioinformatics↗

Mapping cells through time and space with moscot

Single-cell genomics technologies enable multimodal profiling of millions of cells across temporal and spatial dimensions. Experimental limitations prevent the measurement of all-encompassing cellular states in their native temporal dynamics or spatial tissue niche. Optimal transport theory has emerged as a powerful tool to overcome such constraints, enabling the recovery of the original cellular context. However, most algorithmic implementations currently available have not kept up the pace with increasing dataset complexity, so that current methods are unable to incorporate multimodal information or scale to single-cell atlases. Here, we introduce multi-omics single-cell optimal transport (moscot), a general and scalable framework for optimal transport applications in single-cell genomics, supporting multimodality across all applications. We demonstrate moscots ability to efficiently reconstruct developmental trajectories of 1.7 million cells of mouse embryos across 20 time points and identify driver genes for first heart field formation. The moscot formulation can be used to transport cells across spatial dimensions as well: To demonstrate this, we enrich spatial transcriptomics datasets by mapping multimodal information from single-cell profiles in a mouse liver sample, and align multiple coronal sections of the mouse brain. We then present moscot.spatiotemporal, a new approach that leverages gene expression across spatial and temporal dimensions to uncover the spatiotemporal dynamics of mouse embryogenesis. Finally, we disentangle lineage relationships in a novel murine, time-resolved pancreas development dataset using paired measurements of gene expression and chromatin accessibility, finding evidence for a shared ancestry between delta and epsilon cells. Moscot is available as an easy-to-use, open-source python package with extensive documentation at https://moscot-tools.org.

bioinformatics↗

Deep learning-based codon optimization with large-scale synonymous variant datasets enables generalized tunable protein expression

Increasing recombinant protein expression is of broad interest in industrial biotechnology, synthetic biology, and basic research. Codon optimization is an important step in heterologous gene expression that can have dramatic effects on protein expression level. Several codon optimization strategies have been developed to enhance expression, but these are largely based on bulk usage of highly frequent codons in the host genome, and can produce unreliable results. Here, we develop deep contextual language models that learn the codon usage rules from natural protein coding sequences across members of the Enterobacterales order. We then fine-tune these models with over 150,000 functional expression measurements of synonymous coding sequences from three proteins to predict expression in E. coli. We find that our models recapitulate natural context-specific patterns of codon usage and can accurately predict expression levels across synonymous sequences. Finally, we show that expression predictions can generalize across proteins unseen during training, allowing for in silico design of gene sequences for optimal expression. Our approach provides a novel and reliable method for tuning gene expression with many potential applications in biotechnology and biomanufacturing.

synthetic biology↗

Unlocking de novo antibody design with generative artificial intelligence

Generative AI has the potential to redefine the process of therapeutic antibody discovery. In this report, we describe and validate deep generative models for the de novo design of antibodies against human epidermal growth factor receptor (HER2) without additional optimization. The models enabled an efficient workflow that combined in silico design methods with high-throughput experimental techniques to rapidly identify binders from a library of [~]106 heavy chain complementarity-determining region (HCDR) variants. We demonstrated that the workflow achieves binding rates of 10.6% for HCDR3 and 1.8% for HCDR123 designs and is statistically superior to baselines. We further characterized 421 diverse binders using surface plasmon resonance (SPR), finding 71 with low nanomolar affinity similar to the therapeutic anti-HER2 antibody trastuzumab. A selected subset of 11 diverse high-affinity binders were functionally equivalent or superior to trastuzumab, with most demonstrating suitable developability features. We designed one binder with [~]3x higher cell-based potency compared to trastuzumab and another with improved cross-species reactivity1. Our generative AI approach unlocks an accelerated path to designing therapeutic antibodies against diverse targets.

synthetic biology↗

Antibody optimization enabled by artificial intelligence predictions of binding affinity and naturalness

Traditional antibody optimization approaches involve screening a small subset of the available sequence space, often resulting in drug candidates with suboptimal binding affinity, developability or immunogenicity. Based on two distinct antibodies, we demonstrate that deep contextual language models trained on high-throughput affinity data can quantitatively predict binding of unseen antibody sequence variants. These variants span a KD range of three orders of magnitude over a large mutational space. Our models reveal strong epistatic effects, which highlight the need for intelligent screening approaches. In addition, we introduce the modeling of "naturalness", a metric that scores antibody variants for similarity to natural immunoglobulins. We show that naturalness is associated with measures of drug developability and immunogenicity, and that it can be optimized alongside binding affinity using a genetic algorithm. This approach promises to accelerate and improve antibody engineering, and may increase the success rate in developing novel antibody and related drug candidates.

bioinformatics↗