Search bioRxiv⌕ Search

Biology subjects

Csanyi, B.

Publications and source records attributed to Csanyi, B..

3 recordsLinked to original sources

DeltaMut: An Integrative Database of AlphaFold2-Derived Missense Variant Structures

The widespread use of next-generation sequencing has led to a surge in the number of identified variants with uncertain effects on protein function. These variants pose a significant challenge in diagnostics and hinder patient treatment strategies. Numerous variant effect predictors (VEPs) are available to assess variant impact, but they primarily rely on sequence-derived information. The recent development of AlphaFold2 has raised questions about whether information retrieved from wild-type or predicted structures of missense variants can improve the predictive power of these algorithms. While the AlphaFold Protein Structure Database serves as a valuable resource for wild-type protein structures, a large-scale collection of missense variant structures is not available, limiting current efforts to wild-type conformations and a handful of modeled variants. To address this limitation, we developed DeltaMut, a comprehensive database containing over 77,000 protein structures, including 65,000 pathogenic and neutral missense variants. All structural models were generated using ParaFold, a high-performance computing-optimized implementation of AlphaFold2. The large-scale and systematic generation of variant protein structures distinguish DeltaMut as a unique resource for both expansive statistical studies and detailed, case-specific investigations of variant-induced structural changes. Furthermore, the DeltaMut database is freely accessible without registration. HighlightsO_LIDeltaMut is currently the largest database of AlphaFold2-predicted variant structures. C_LIO_LIContains 77,713 structures covering 12,101 wild-type and 65,612 variant proteins. C_LIO_LI70.6% of predicted structures have high or very high confidence (pLDDT [≥] 70). C_LIO_LIFreely accessible web server with visualization and download of variant models. C_LI

bioinformatics↗

Assessing the impact of parental linear gene normalization on the performance of statistical models for circular RNA differential expression analysis

BackgroundCircular RNAs (circRNAs) emerged as promising non-invasive cancer biomarkers due to their stability, abundance in body fluids, and regulatory potential. However, circRNA differential expression analysis (DEA) remains challenging, largely owing to lack of consensus on important preprocessing strategies such as filtering and normalization. While well-established bulk RNA-sequencing frameworks are commonly applied to circRNA data, newer approaches such as CIRI-DE (part of CIRI3 suite) integrate both linear and circular transcript information to improve detection. Despite developments, an assessment of these integrative strategies is lacking, and the critical impact of filtering on DEA model performance has not been comprehensively evaluated. ResultsIn this study, we evaluated the impact of multiple normalization and filtering strategies on circRNA DEA using five experimental datasets, including two in-house blood platelet sets and semi-parametric simulated in silico datasets. Our results emphasize the importance of selecting an appropriate filtering threshold, as overly lenient filtering substantially reduced model performance across datasets. We found edgeRs filterByExpr() strategy particularly effective in handling zero counts in circRNA data, while also generating the most reliable results across most datasets. Furthermore, by incorporating linear and circular information as described in CIRI-DE, most methods identified a higher number of differentially expressed (DE) circRNAs compared to circular counts alone. Notably, circRNAs identified by both CIRI-DE and the modified bulk RNA-sequencing pipelines showed substantial overlap. ConclusionOur findings demonstrate that automated filtering combined with linear-aware normalization significantly enhances the sensitivity and reproducibility of circRNA DEA, providing a standardized framework for more reliable biomarker discovery in transcriptomic research.

bioinformatics↗

Comprehensive Bulk and Single-Cell RNA Sequencing Uncovers Senescence-Associated Biomarkers in Therapeutic Mesenchymal Stem Cells

BackgroundMesenchymal stem cells (MSCs) hold great promise in cell therapy, but their effectiveness declines with repeated cell divisions due to senescence. Canines, sharing aging characteristics with humans, serve as a valuable model to study this process in a translational context. MethodsIn the present study, we performed an in-depth characterization of senescence in canine MSCs using a combination of morphological, molecular, and transcriptomic analyses. Early (P2) and late-passage (P6) canine MSCs were characterized using a combination of senescence-associated {beta}-galactosidase staining, cell cycle profiling, and both bulk and single-cell RNA sequencing to capture global transcriptional changes. ResultsBy employing a passage-based in vitro approach, the present study demonstrates that late-passage cells (P6) compared to early-passage cells (P2) exhibit hallmark features of senescence, including morphological alterations, elevated SA-{beta}-galactosidase activity, and considerable transcriptional changes. These changes were represented by significant upregulation of established senescence marker genes, alongside potential novel candidates and downregulation of genes associated with cell cycle progression and proliferation. Moreover, single-cell RNA sequencing uncovered heterogeneous distribution of senescent subpopulations, upregulation of SASP-related genes and reduced proliferation markers. ConclusionsOur findings demonstrate that combining classical markers with bulk and single-cell RNA sequencing facilitates senescent cell identification while improving quality control for clinical MSC samples.

cell biology↗