Search bioRxiv⌕ Search

Biology subjects

Enderti, A.

Publications and source records attributed to Enderti, A..

2 recordsLinked to original sources

Expert-Guided Supervised Annotation of Erythroid Differentiation in Single-Cell RNA-seq

Accurate annotation of intermediate cell states remains a major challenge in single-cell RNA sequencing (scRNA-seq), particularly in continuous differentiation systems such as erythropoiesis. Existing reference-based methods often lack the resolution required to distinguish early and transitional erythroid progenitors and may generalise poorly across datasets and modalities. Here, we present a supervised framework for erythroid lineage annotation based on expert-curated training data that integrates bulk and single-cell transcriptomic information. Starting from a human bone marrow scRNA-seq atlas, we refined erythroid annotations by introducing previously unresolved progenitor stages, including burst-forming unit-erythroid (BFU-E), colony-forming unit-erythroid (CFU-E), and pro-erythroblast (ProE), guided by canonical marker genes and bulk RNA-seq references. We trained and benchmarked four classical machine learning models and identified LightGBM as the best-performing approach, achieving a validation macro F1-score of 0.821 and balanced accuracy of 0.826. On a held-out test set, the model showed strong performance across most erythroid stages, with errors largely confined to adjacent differentiation states. The classifier was further transferred to independent bulk RNA-seq samples and an external bone marrow scRNA-seq dataset, where it recovered expected erythroid progression and refined coarse-grained annotations into higher-resolution cell states. Together, these results show that expert-curated supervised learning can improve erythroid cell state annotation in scRNA-seq and provide a practical framework for studying differentiation hierarchies in settings where finely resolved public references are limited.

bioinformatics↗

Enformer-Based Phylogenetic Tree Reconstruction

Enformer is a deep learning model trained on human and mouse genomes to predict regulatory activity from 196,608 bp DNA windows. Its trunk embeddings capture long-range cis-regulatory interactions, but whether this signal generalises across the tree of life has not been assessed. We embed universal single-copy orthologous groups (OGs) from OrthoDB v12 across three taxonomic scales and evaluate reconstructed trees against TimeTree5 using Mantel r and Normalised Robinson-Foulds (NRF). On 702 OGs across 34 Primate species ([≤] 74 Mya), the consensus tree achieves Mantel r = 0.902 and NRF = 0.481, correctly recovering major clades. A key finding is that flanking regulatory context -- not the gene locus itself -- carries the phylogenetic signal: restricting pooling to central 448 bins collapses Mantel r to 0.355. Applying the same fixed configuration to Vertebrates ([≤] 450 Mya, 83 OGs, 150 species) and Plants ([≤] 1,500 Mya, 92 OGs, 40 species) yields consensus Mantel r of 0.752 and 0.803 respectively, with NRF worsening monotonically across tiers. Distance-ordering fidelity degrades smoothly with evolutionary distance while topological accuracy declines steadily, with no sharp taxonomic boundary. These results show that an unmodified regulatory deep learning model encodes robust phylogenetic signal well beyond its training distribution, reaching across 1,500 million years of divergence.

bioinformatics↗