Search bioRxiv⌕ Search

Biology subjects

Baragli, M.

Publications and source records attributed to Baragli, M..

2 recordsLinked to original sources

Benchmarking DNA Foundation Models for zero-shot variant effect prediction: the role of context, training, and architecture

In this study, we systematically evaluate the performance of several DNA foundation models (NT, DNABERT, and HyenaDNA) in predicting the functional impact of genetic variants using Zero-shot scoring, a method that does not require task-specific fine-tuning. We assess the models sensitivity to sequence alterations introduced by Single Nucleotide Variants (SNVs), comparing their ability to capture both local and extended contextual effects. Using pathogenic, benign, and uncertain SNVs from ClinVar, we show that large multi-species NT models outperform other architectures in detecting functional consequences, not only at the mutation site but also in adjacent regions. These models exhibit superior discriminative power across variant categories, especially when aggregating Zero-shot scores over multiple surrounding tokens. Conversely, models trained solely on human sequences, such as DNABERT and HyenaDNA, show limited contextual awareness and reduced ability to differentiate variant effects. Our findings highlight the critical importance of model size, training objective, and training data diversity in shaping model performance. Furthermore, we discuss current limitations in modeling long-range dependencies in genomic sequences and suggest that innovations in transformer architectures, such as sparse attention or memory-augmented models, may provide viable paths toward scalable, genome-wide variant effect prediction.

genomics↗

PoreMeth2: decoding the evolution of methylome alterations with Nanopore sequencing

In epigenetic analysis, identifying differentially methylated regions (DMRs) typically involves detecting groups of consecutive CpGs that show significant changes in their average methylation levels. However, the methylation state of a genomic region can also be characterized by a mixture of patterns (epialleles) with variable frequencies, and the relative proportions of such patterns can provide insights into its mechanisms of formation. Traditional methods based on bisulfite conversion and NGS, due to the read size (150 bp), allow epiallele frequency analysis only in high-CpG-density regions, limiting differential methylation studies to just 50% of the human methylome. Nanopore sequencing, with its long reads, enables the analysis of epiallele frequency across both high- and low-CpG-density regions. We introduce a novel computational approach, PoreMeth2, an R library that integrates epiallelic diversity and methylation frequency changes from Nanopore data to identify DMRs, assess their formation mechanisms, and annotate them to genic and regulatory elements. We applied PoreMeth2 to cancer and glial cell datasets, demonstrating its ability to distinguish epigenomic changes with a strong effect on gene expression from those with a weaker impact on transcriptional activity. PoreMeth2 is publicly available at https://github.com/Lab-CoMBINE/PoreMeth2.

bioinformatics↗