Search bioRxiv⌕ Search

Biology subjects

Ruggeri, L.

Publications and source records attributed to Ruggeri, L..

3 recordsLinked to original sources

Improving DNA Modeling with WaveDNA: Enhancing Speed, Generalizability, and Interpretability through Wavelet Transformation

Transcription factors (TFs) regulate gene expression by binding to short, specific DNA sequences, known as transcription factor binding sites (TFBSs). Accurate identification of TFBSs is fundamental for understanding transcriptional regulation. By leveraging their ability to capture complex non-linear patterns and hierarchical dependencies underlying TF-DNA binding deep learning (DL) has emerged as the state-of-the-art approach for modeling and identifying TFBSs. However, current models often require extensive pretraining, involve large parameter sets, and offer limited interpretability. To address these limitations, we introduce WaveDNA, a lightweight and interpretable DL framework that encode DNA sequences into two-dimensional representations using wavelet transforms. This approach enables the use of convolutional neural networks pretrained on images, facilitating efficient transfer learning without requiring large-scale genomic data pretraining. Across diverse ENCODE ChIP-seq datasets spanning different TFs, WaveDNA achieves predictive accuracy comparable to state-of-the-art DL models while using approximately fivefold fewer parameters and substantially less computational resources. Moreover, representing DNA sequences as images allows the direct application of established computer vision interpretability techniques to visualize the learned binding patterns. Together, these results demonstrate that WaveDNA offers a scalable, computationally efficient, and interpretable alternative for modeling TF-DNA interactions.

bioinformatics↗

Benchmarking PWM and SVM-based Models for Transcription Factor Binding Site Prediction: A Comparative Analysis on Synthetic and Biological Data

Transcription Factors (TFs) are essential regulatory proteins that control the cellular transcriptional states by binding to specific DNA sequences known as Transcription Factor Binding Sites (TFBSs) or motifs. Accurate TFBS identification is crucial for unraveling regulatory mechanisms driving cellular dynamics. Over the years, various computational approaches have been developed to model TFBSs, with Position Weight Matrices (PWMs) being one of the most widely adopted methods. PWMs provide a probabilistic framework by representing nucleotide frequencies at every position within the binding site. While effective and interpretable, PWMs face significant limitations, such as their inability to capture positional dependencies or model complex interactions. To address these, advanced methods, such as Support Vector Machine (SVM)-based models, have been introduced. Leveraging human ChIP-seq data from ENCODE, this study systematically benchmarks the predictive performance of PWM and SVM-based models across different scenarios. We evaluate the impact of key factors such as training dataset size, sequence length, and kernel functions (for SVMs) on models performance. Additionally, we explore the impact of synthetic versus real biological background data during model training. Our analysis highlights strengths and limitations of both PWM and SVM-based approaches under different conditions, providing practical guidance for selecting and tailoring models to specific biological datasets. To complement our analysis, we present a comprehensive database of pretrained SVM models for TFBS detection, trained on human ChIP-seq data from diverse cell lines and tissues. This resource aims to facilitate broader adoption of SVM-based methods in TFBS prediction and enhance their practical utility in regulatory genomics research.

bioinformatics↗

A systemically delivered AAV-CFTR gene therapy for cystic fibrosis

Cystic fibrosis (CF) is the most common monogenic lung disease and results from mutations in the Cystic Fibrosis Transmembrane Conductance Regulator (CFTR). There have been over 2000 variants identified in patients that result in loss of function of the CFTR protein leading to systemic disease and respiratory failure in adolescence. While some variants encode proteins with residual activity that can be corrected or potentiated by CFTR modulators, at least 10% of CF individuals cannot tolerate the modulators or have nonsense mutations which fail to make any protein. For all people with CF, a mutation agnostic gene replacement strategy could provide a cure for CF lung disease. Here, we propose using a systemic route of administration to deliver a functional CFTR minigene cargo with a lung tropic AAV capsid. This would serve to reach multiple organs, most importantly the lung epithelium, and would provide a functional CFTR transgene that could be expressed in any cell type with a ubiquitous promoter. To achieve this, we generated the smallest CFTR minigene tested in an AAV delivery to date. We demonstrate its expression and function following transfection in cell-based assays and restoration of function in primary CF airway cells after viral delivery. Furthermore, we identify an AAV capsid that can transduce alveolar and airway epithelium with systemic delivery in non-human primates. These data provide tools for delivering a functional CFTR minigene that fits within the packaging capacity of an AAV and demonstrate lung transduction with an AAV following systemic delivery in a large animal model. This strategy first and foremost can reach target airway cells by circumventing the strong mucosal barrier in CF airways but may also provide a method by which to restore CFTR function in additional CF affected organs.

genetics↗