Search bioRxiv⌕ Search

Biology subjects

Stranieri, N.

Publications and source records attributed to Stranieri, N..

4 recordsLinked to original sources

CREP: Cis-Regulatory Element Predictor Based on Fine-Tuned Enformer

A substantial fraction of disease-associated genetic variants reside in non-coding regions of the genome, where they act by perturbing cis-regulatory elements (CREs) such as enhancers, promoters, and insulators. While recent sequence-based deep learning models, such as Enformer, accurately predict continuous epigenomic signals from DNA sequence, they do not directly provide discrete and interpretable CRE annotations. Here, we present CREP (Cis-Regulatory Element Predictor), a fine-tuned version of Enformer trained to predict regulatory element identity from sequence using REgulamentary-derived annotations across multiple human cell-types. Through a controlled experimental framework, we show that incorporating diverse cell-types improves model performance. CREP leverages cell-type-specific training data to learn regulatory representations while producing a unified prediction of CRE identity from sequence. This is demonstrated by the Vanuatu SNP, a non-coding variant that creates a de novo erythroid regulatory element, which is correctly detected only when erythroid data are included during training. Error analysis further reveals that apparent misclassifications between enhancers and promoters reflect their shared regulatory architecture, supporting the view of CREs as a functional continuum rather than strictly discrete classes. Together, these results demonstrate that CREP enables interpretable prediction of regulatory element identity from sequence and provides a framework for the functional interpretation of non-coding genetic variation.

bioinformatics↗

Expert-Guided Supervised Annotation of Erythroid Differentiation in Single-Cell RNA-seq

Accurate annotation of intermediate cell states remains a major challenge in single-cell RNA sequencing (scRNA-seq), particularly in continuous differentiation systems such as erythropoiesis. Existing reference-based methods often lack the resolution required to distinguish early and transitional erythroid progenitors and may generalise poorly across datasets and modalities. Here, we present a supervised framework for erythroid lineage annotation based on expert-curated training data that integrates bulk and single-cell transcriptomic information. Starting from a human bone marrow scRNA-seq atlas, we refined erythroid annotations by introducing previously unresolved progenitor stages, including burst-forming unit-erythroid (BFU-E), colony-forming unit-erythroid (CFU-E), and pro-erythroblast (ProE), guided by canonical marker genes and bulk RNA-seq references. We trained and benchmarked four classical machine learning models and identified LightGBM as the best-performing approach, achieving a validation macro F1-score of 0.821 and balanced accuracy of 0.826. On a held-out test set, the model showed strong performance across most erythroid stages, with errors largely confined to adjacent differentiation states. The classifier was further transferred to independent bulk RNA-seq samples and an external bone marrow scRNA-seq dataset, where it recovered expected erythroid progression and refined coarse-grained annotations into higher-resolution cell states. Together, these results show that expert-curated supervised learning can improve erythroid cell state annotation in scRNA-seq and provide a practical framework for studying differentiation hierarchies in settings where finely resolved public references are limited.

bioinformatics↗

Enformer-Based Phylogenetic Tree Reconstruction

Enformer is a deep learning model trained on human and mouse genomes to predict regulatory activity from 196,608 bp DNA windows. Its trunk embeddings capture long-range cis-regulatory interactions, but whether this signal generalises across the tree of life has not been assessed. We embed universal single-copy orthologous groups (OGs) from OrthoDB v12 across three taxonomic scales and evaluate reconstructed trees against TimeTree5 using Mantel r and Normalised Robinson-Foulds (NRF). On 702 OGs across 34 Primate species ([≤] 74 Mya), the consensus tree achieves Mantel r = 0.902 and NRF = 0.481, correctly recovering major clades. A key finding is that flanking regulatory context -- not the gene locus itself -- carries the phylogenetic signal: restricting pooling to central 448 bins collapses Mantel r to 0.355. Applying the same fixed configuration to Vertebrates ([≤] 450 Mya, 83 OGs, 150 species) and Plants ([≤] 1,500 Mya, 92 OGs, 40 species) yields consensus Mantel r of 0.752 and 0.803 respectively, with NRF worsening monotonically across tiers. Distance-ordering fidelity degrades smoothly with evolutionary distance while topological accuracy declines steadily, with no sharp taxonomic boundary. These results show that an unmodified regulatory deep learning model encodes robust phylogenetic signal well beyond its training distribution, reaching across 1,500 million years of divergence.

bioinformatics↗

REnformer, a single-cell ATAC-seq predicting model to investigate open chromatin sites

Genome regulatory elements are fundamental to cellular identity and cell type specific gene expression. Understanding how the underlying genetic code is differentially utilised by different cell types is central to understanding human health and disease. To better understand how DNA encodes genome regulatory elements such as promoters, enhancers, and boundary elements, we leverage the Enformer gene expression and epigenetic prediction model. We used transfer learning with high quality single cell ATAC datasets to develop REnformer, a model to predict chromatin accessibility. By introducing a bench-mark for comparing performances against Enformer model, REnformer significantly outperformed Enformer in terms of higher prediction outcomes and lower error rates in all extensive analyses shown; introducing these benchmarks allowed us, and possible future works, to fairly compare such models. We further tested REnformer by predicting the effects of a well characterised -thalassemia variant and found that the prediction aligned with the observed change in genome regulatory element, previously validated. We conclude that REnformer is and can be a state-of-the-art tool to predict cell type specific regulatory elements, and interrogate the effect of genome variation in health and disease.

bioinformatics↗