Search bioRxiv⌕ Search

Biology subjects

Mahieu, L.

Publications and source records attributed to Mahieu, L..

4 recordsLinked to original sources

GAME: Genomic API for Model Evaluation

The rapid expansion of genomics datasets and the application of machine learning has produced sequence-to-activity genomics models with ever-expanding capabilities. However, benchmarking these models on practical applications has been challenging because individual projects evaluate their models in ad hoc ways, and there is substantial heterogeneity of both model architectures and benchmarking tasks. To address this challenge, we have created GAME, a system for large-scale, community-led standardized model benchmarking on user-defined evaluation tasks. We borrow concepts from the Application Programming Interface (API) paradigm to allow for seamless communication between pre-trained models and benchmarking tasks, ensuring consistent evaluation protocols. Because all models and benchmarks are inherently compatible in this framework, the continual addition of new models and new benchmarks is easy. We also developed a Matcher module powered by a large language model (LLM) to automate ambiguous task alignment between benchmarks and models. Containerization of these modules enhances reproducibility and facilitates the deployment of models and benchmarks across computing platforms. By focusing on predicting underlying biochemical phenomena (e.g. gene expression, open chromatin, DNA binding), we ensure that tasks remain technology-independent. We provide examples of benchmarks and models implementing this framework, and anticipate that the community will contribute their own, leading to an ever-expanding and evolving set of models and evaluation tasks. This resource will accelerate genomics research by illuminating the best models for a given task, motivating novel functional genomic benchmarks, and providing a more nuanced understanding of model abilities.

bioinformatics↗

Decoding cnidarian cell type gene regulation

Animal cell types are defined by differential access to genomic information, a process orchestrated by the combinatorial activity of transcription factors that bind to cis-regulatory elements (CREs) to control gene expression. However, the regulatory logic and specific gene networks that define cell identities remain poorly resolved across the animal tree of life. As early-branching metazoans, cnidarians can offer insights into the early evolution of cell type-specific genome regulation. Here, we profiled chromatin accessibility in 60,000 cells from whole adults and gastrula-stage embryos of the sea anemone Nematostella vectensis. We identified 112,728 CREs and quantified their activity across cell types, revealing pervasive combinatorial enhancer usage and distinct promoter architectures. To decode the underlying regulatory grammar, we trained sequence-based models predicting CRE accessibility and used these models to infer ontogenetic relationships among cell types. By integrating sequence motifs, transcription factor expression, and CRE accessibility, we systematically reconstructed the gene regulatory networks that define cnidarian cell types. Our results reveal the regulatory complexity underlying cell differentiation in a morphologically simple animal and highlight conserved principles in animal gene regulation. This work provides a foundation for comparative regulatory genomics to understand the evolutionary emergence of animal cell type diversity.

genomics↗

HyDrop v2: Scalable atlas construction for training sequence-to-function models

Deciphering cis-regulatory logic underlying cell type identity is a fundamental question in biology. Single-cell chromatin accessibility (scATAC-seq) data has enabled training of sequence-to-function deep learning models allowing decoding of enhancer logic and design of synthetic enhancers. Training such models requires large amounts of high-quality training data across species, organs, development, aging, and disease. To facilitate the cost-effective generation of large scATAC-seq atlases for model training, we developed a new version of the open-source microfluidic system HyDrop with increased sensitivity and scale: HyDrop v2. We generated HyDrop v2 atlases for the mouse cortex and Drosophila embryo development and compared them to atlases generated on commercial platforms. HyDrop v2 data integrates seamlessly with commercially available chromatin accessibility methods (10x Genomics). Differentially accessible regions and motif enrichment across cell types are equivalent between HyDrop-v2 and 10x atlases. Sequence-to-function models trained on either atlas are comparable as well in terms of enhancer predictions, sequence explainability, and transcription factor footprinting. By offering accessible data generation, enhancer models trained on HyDrop-v2 and mixed atlases can contribute to unraveling cell-type specific regulatory elements in health and disease.

bioinformatics↗

CREsted: modeling genomic and synthetic cell type-specific enhancers across tissues and species

Sequence-based deep learning models have become the state of the art for the analysis of the genomic regulatory code. Particularly for transcriptional enhancers, deep learning models excel at deciphering sequence features and grammar that underlie their spatiotemporal activity. To enable end-to-end enhancer modeling and design, we developed a software and modeling package, called CREsted. It combines preprocessing starting from single-cell ATAC-seq data; modeling with a choice of several architectures for training classification and regression models on either topics or pseudobulk peak heights; sequence design using multiple strategies; and downstream analysis through a collection of tools to locate transcription factor (TF) binding sites, infer the effect of a TF (activating or repressing) on enhancer accessibility, decipher enhancer grammar, and score gene loci. We demonstrate CREsted using a mouse cortex model that we validate using the BICCN collection of in vivo validated mouse brain enhancers. Classical enhancers in immune cells, including the IFNB1 enhanceosome are revisited using a PBMC model, and we assess the accuracy of TF binding site predictions with ChIP-seq. Additionally, we use CREsted to compare mesenchymal-like cancer cell states between tumor types; and we investigate different fine-tuning strategies of Borzoi within CREsted, comparing their performance and explainability with CREsted models trained from scratch. Finally, we train a CREsted model on a scATAC-seq atlas of zebrafish development and use this to design and in vivo validate cell type-specific synthetic enhancers in three tissues. For varying datasets, we demonstrate that CREsted facilitates efficient training and analyses, enabling scrutinization of the enhancer logic and design of synthetic enhancers across tissues and species. CREsted is available at https://crested.readthedocs.io.

genomics↗