Search bioRxiv⌕ Search

Biology subjects

Pearce, J.

Publications and source records attributed to Pearce, J..

3 recordsLinked to original sources

Species-specific small models for cell type classification approach the performance of large single cell foundation models

Accurate cross-species cell type classification remains a key evaluation task in single-cell transcriptomics. Recent foundation models trained on millions of single-cell profiles demonstrate great in-distribution and out-of-distribution performance on this task, but their large parameter counts and substantial computational costs limit accessibility and interpretability. Here, we introduce CytoType, a simple and interpretable model for cell type classification that leverages pre-trained ESM-2 protein embeddings of protein-coding transcripts. By learning linear, cell-type-specific weights over transcript embeddings, without relying on gene count information, CytoType achieves F1 scores comparable to or exceeding those of large-scale transformer-based models. We further developed ESM-Cell Embedding (ESM-CE), an even simpler variant that only averages ESM-2 embeddings across expressed genes, which also performs competitively against foundation models. Both CytoType and ESM-CE are trained on species-specific data, maintaining high accuracy when classifying cell types with orders of magnitude fewer parameters compared to larger foundation models. For example, for human tissues, the average performance gap between CytoType and the best foundation model was 0.053 F1 points while CytoType uses 10,000x fewer trainable parameters. Additionally, we quantified the contribution of ESM-2 embeddings to cell type classification tasks and demonstrated a three fold reduction in the performance gap between CytoType and the best foundation model for nine species. Finally, we show that CytoTypes learned gene weights are biologically interpretable.

cell biology↗

Epigenetic priming of embryonic enhancer elements coordinates developmental gene networks

Embryonic development requires the accurate spatiotemporal execution of cell lineage-specific gene expression programs, which are controlled by transcriptional enhancers. Developmental enhancers adopt a primed chromatin state prior to their activation; however how this primed enhancer state is established, maintained, and how it affects the regulation of developmental gene networks remains poorly understood. Here, we use comparative multi-omic analyses of human and mouse early embryonic development to identify subsets of post-gastrulation lineage-specific enhancers which are epigenetically primed ahead of their activation, marked by the histone modification H3K4me1 within the epiblast. We show that epigenetic priming occurs at lineage-specific enhancers for all three germ layers, and that epigenetic priming of enhancers confers lineage-specific regulation of key developmental gene networks. Surprisingly in some cases, lineage-specific enhancers are epigenetically marked already in the zygote, weeks before their activation during lineage specification. Moreover, we outline a generalisable strategy to use naturally occurring human genetic variation to delineate important sequence determinants of primed enhancer function. Our findings identify an evolutionarily conserved program of enhancer priming and begin to dissect the temporal dynamics and mechanisms of its establishment and maintenance during early mammalian development.

developmental biology↗

A comprehensive evaluation of taxonomic classifiers in marine vertebrate eDNA studies

Environmental DNA (eDNA) metabarcoding is a widely used tool for surveying marine vertebrate biodiversity. To this end, many computational tools have been released and a plethora of bioinformatic approaches are used for eDNA-based community composition analysis. Simulation studies and careful evaluation of taxonomic classifiers are essential to establish reliable benchmarks to improve accuracy and reproducibility of eDNA-based findings. Here we present a comprehensive evaluation of nine taxonomic classifiers exploring three widely used mitochondrial markers (12S rDNA, 16S rDNA, and COI) in Australian marine vertebrates. Curated reference databases and exclusion database tests were used to simulate diverse species compositions, including three positive control and two negative control datasets. Using these simulated datasets, we were able to identify between 19% to 85% of marine vertebrate species using mitochondrial markers. We show that MMSeqs2 and Metabuli generally outperform BLAST with 10% and 11% higher F1 scores for 12S and 16S rDNA markers, respectively, and that Naive Bayes Classifiers such as Mothur outperform sequence-based classifiers except MMSeqs2 for COI markers by 11%. Database exclusion tests reveal that MMSeqs2 and BLAST are less susceptible to false positives compared to Kraken2 with default parameters. Based on these findings, we recommend that MMSeqs2 is used for taxonomic classification of marine vertebrates given its ability to improve species-level assignments while reducing the number of false positives. Our work contributes to the establishment of best practices in eDNA-based biodiversity analysis to ultimately increase the reliability of this monitoring tool in the context of marine vertebrate conservation.

ecology↗