Search bioRxiv⌕ Search

Biology subjects

Daher, A.

Publications and source records attributed to Daher, A..

4 recordsLinked to original sources

A Robust Machine Learning Framework for Keloid Biomarker Discovery Beyond Differential Expression

Keloids are fibroproliferative skin disorders arising following dermal injury that extend beyond the original wound margins. Their pathogenesis remains poorly understood, and current treatments are associated with high recurrence rates. Identifying transcriptomic biomarkers that distinguish keloids from other skin and scar phenotypes may provide insight into disease mechanisms and facilitate the development of targeted therapeutic approaches. However, previous transcriptomic studies have often been limited by small sample sizes, pairwise comparisons between tissue classes, heterogeneous data-integration strategies, and a reliance on conventional differential gene expression (DGE) analysis. Here, we employed a multi-stage machine learning (ML) workflow for robust keloid biomarker discovery using transcriptomic datasets derived from both bulk RNA sequencing and single-cell RNA sequencing (scRNA-seq). We assembled and harmonized, to the best of our knowledge, the largest curated cross-study keloid transcriptomic cohort currently available, comprising 81 samples from 13 independent studies spanning four clinically relevant tissue classes: normal skin, normotrophic scar, hypertrophic scar, and keloid scar. Through study-aware cross-validation, feature selection, partition-stability analysis, and bootstrap validation across multiple ML classifiers, we identified a panel of eight highly consistent biomarkers capable of distinguishing keloid from non-keloid samples. These biomarkers were associated with dysregulation of extracellular matrix homeostasis, fibrosis-resolution pathways, vascular remodelling, and metabolic reprogramming. Comparison with conventional DGE analysis demonstrated substantial agreement while also highlighting important differences between the two approaches. In particular, FASN was consistently identified by the ML workflow as an upregulated discriminatory biomarker despite exhibiting weak, non-significant differential expression in the DGE analysis. Cell-type-specific analysis further supported this finding, revealing significant FASN upregulation in fibroblast and vascular endothelial populations. These results demonstrate that ML and DGE capture complementary aspects of transcriptomic variation. This study provides a robust strategy for cross-study transcriptomic biomarker discovery and identifies candidate genes and pathways for future mechanistic and therapeutic investigation in keloids. 1 Author SummaryKeloids are abnormal scars that continue to grow beyond the original wound and can be difficult to treat because they frequently recur after therapy. Although many studies have investigated the biology of keloids, the molecular mechanisms that distinguish them from other scar types remain incompletely understood. Identifying biomarkers involved in keloid formation may help inform improved treatment strategies. Previous transcriptomic studies have often been limited by small sample sizes and inconsistent analytical approaches. In this study, we combined gene-expression data from multiple independent studies to create, to the best of our knowledge, the largest cross-study transcriptomic collection available for keloid analysis. We then applied several machine learning approaches to identify genes that consistently distinguished keloids from other skin and scar phenotypes. The identified biomarkers were associated with extracellular matrix remodeling, fibrosis, vascular function, and cellular metabolism. One gene involved in fatty-acid synthesis, FASN, was repeatedly identified by the machine learning analyses despite being overlooked by conventional gene-expression methods. Additional single-cell analyses confirmed elevated FASN expression in specific cell populations within keloid tissue. More broadly, this work provides a strategy for discovering robust biomarkers from heterogeneous biological datasets and identifies molecular targets for future studies of keloid disease.

bioinformatics↗

Recognition and Resolution of KRAS 5'UTR RNA G-Quadruplexes by hnRNPA1

The KRAS oncogene, central to cellular signaling via MAPK and PI3K-AKT pathways, is a notorious cancer driver frequently activated in pancreatic, colorectal, and lung carcinomas. Regulation of human KRAS oncogene expression is important due to its capital role in cell growth, proliferation, and survival. Misregulation of its expression contributes directly to the development and progression of multiple types of cancer. In previous studies, the role of G-quadruplexes elements in both the promoter and 5 UTR regions have shown to play important roles in KRAS expression, particularly when these G4s elements interact with regulatory protein hnRNPA1. In this study, we reveal that KRAS expression is also modulated at the post-transcriptional level through the formation of RNA G-quadruplexes (rG4s) situated at the 5 untranslated region (5UTR) of the mRNA. Biophysical and binding studies were carried out to probe the interaction. Through isothermal titration calorimetry (ITC), we quantified a strong binding affinity between the UP1 domain of hnRNPA1 and short-nucleotide RNA segments capable of adopting different G-quadruplex fold. The binding interaction is characterized by a favorable Gibbs free energy change in the range of {Delta}G {approx} -32 to -34 kJ/mol, suggesting a specific and energetically favorable association. One-dimensional and two-dimensional 1H-15N HSQC NMR spectroscopy revealed pronounced chemical shift changes in residues of both RNA recognition motifs (RRMs) of UP1, signifying direct contact with the rG4 structure.

biophysics↗

A Computational Pipeline for Physiologically Informed Calibration of Ligand Reaction-Diffusion Models Using High-Throughput Sequencing

All physiological processes fundamentally rely on continuous cellular cross-talk to maintain organization and ensure proper function. Among the various modes of cellular communication, ligand-mediated chemical signaling, in which a ligand is secreted by one cell, diffuses through the extracellular environment, and binds to a receptor on another (or the same) cell to elicit a downstream response, is arguably the most ubiquitous and foundational. Given its importance, numerous mathematical models have been developed to describe this reaction-diffusion mechanism, capturing ligand secretion, diffusion, decay, and binding under both normal and pathological conditions. However, parameter calibration for such models often lags behind model development. This is due to limited data that faithfully represent the biological microenvironment, as well as due to the absence of a robust, rigorous framework to integrate available data into the mathematical model. To address this gap, we propose that transcriptomics (gene expression) data, namely the combination of single-cell RNA sequencing and spatial transcriptomics, provide a rich, increasingly abundant, and underutilized source of information that can be used to calibrate the parameters of the cellular reaction-diffusion models at the larger mesoscopic scale. To this end, we develop a computational pipeline that leverages these data to extract parameter values for reaction-diffusion models, and illustrate its application through two human wound-healing case studies. Using open-source transcriptomics data, we calibrate the reaction-diffusion model parameters of the isoforms of Transforming Growth Factor Beta (TGF{beta}), a signaling molecule central to tissue repair and development as well as to pathological processes such as cancer and fibrosis. Our pipeline integrates traditional numerical (finite volume) solvers for the ligand concentration fields with bioinformatics, machine learning, and Bayesian inference methods, combining existing and novel computational tools into a single framework for a physiologically informed, data-driven parameter calibration process. The pipeline is modular, allowing easy extension or adjustment depending on user needs. Overall, this framework facilitates rigorous model calibration, an essential step toward ensuring that mathematical models have meaningful research and potential translational utility.

systems biology↗

The sequestration of miR-642a-3p by a complex formed by HIV-1 Gag and human Dicer increases AFF4 expression and viral production

Micro (mi)RNAs are critical regulators of gene expression in human cells, the functions of which can be affected during viral replication. Here, we show that the human immunodeficiency virus type 1 (HIV-1) structural precursor Gag protein interacts with the miRNA processing enzyme Dicer. RNA immunoprecipitation and sequencing experiments show that Gag modifies the retention of a specific miRNA subset without affecting Dicers pre- miRNA processing activity. Among the retained miRNAs, miR-642a-3p shows an enhanced occupancy on Dicer in the presence of Gag and is predicted to target AFF4 mRNA, which encodes an essential scaffold protein for HIV-1 transcriptional elongation. miR-642a-3p gain- or loss-of-function negatively or positively regulates AFF4 protein expression at mRNA and protein levels with concomitant modulations of HIV-1 production, consistent with an antiviral activity. By sequestering miR-642a-3p with Dicer, Gag enhances AFF4 expression and HIV- 1 production without affecting miR-642a-3p levels. These results identify miR-642a-3p as a strong suppressor of HIV-1 replication and uncover a novel mechanism by which a viral structural protein directly disrupts an miRNA function for the benefit of its own replication. IMPORTANCEVirus-host relationships occur at different levels and the human immunodeficiency virus type 1 (HIV-1) can modify the expression of microRNAs in different cells. Here, we identify a virus- host interaction between the HIV-1 structural protein Gag and the miRNA-processing enzyme Dicer. Gag does not affect the microRNA processing function of Dicer but affects the functionality of a subset of microRNAs that are enriched on the Dicer-Gag complex compared to on Dicer alone. We show that miR-642a-3p, the most enriched microRNA on the Dicer- Gag complex targets and degrades AFF4 mRNA coding for a protein from the super transcription elongation complex, essential for HIV-1 and cellular transcription. Interestingly, the silencing capacity by miR-642a-3p is hindered by Gag and heightened in its absence, consequently affecting HIV-1 transcription. These findings unveil a new paradigm that a microRNA function rather than its abundance can be affected by a viral protein through its enhanced retention on Dicer.

microbiology↗