Search bioRxiv⌕ Search

Biology subjects

HARUN-OR-ROSHID, M.

Publications and source records attributed to HARUN-OR-ROSHID, M..

2 recordsLinked to original sources

Meta-PseU: A Meta-Classifier for Robust Prediction of RNA Pseudouridine Modification Sites from Long Sequences

Pseudouridine ({Psi}) represents one of the most abundant and evolutionarily conserved RNA modifications. {Psi} provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of {Psi} sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, current machine-learning and deep-learning predictors suffer from limitations such as small datasets and limited generalizability. To overcome these issues, we have constructed new long-sequence datasets derived from RMBase 3.0 and developed Meta-PseU, a logistic regression-based meta-classifier that stacks multiple single-feature or baseline classifiers across three species of human, mouse, and yeast. Meta-PseU substantially reduced the performance gaps between training and independent test datasets, presenting superior generalization. Meta-PseU substantially outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. This work offers a new framework for robust {Psi}-site identification by using long sequences. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU.

bioinformatics↗

GDA-Pred: Generative AI-Driven Data Augmentation for Improved Prediction of IL-6 and IL-13 Inducing Peptides

The identification of interleukin-6 (IL-6) and interleukin-13 (IL-13) inducing peptides is crucial for accelerating drug discovery targeting cancer, immune disorders, and infectious diseases. However, experimental screening methods remain time-consuming and costly. To address these limitations, various machine learning and deep learning models have been developed, yet their performance is still constrained by the limited availability of experimentally validated data. In this study, we propose a generative AI-driven data augmentation (GDA) framework and predictor, GDA-Pred, to improve the prediction performance of state-of-the-art (SOTA) classifiers for identifying IL-6 and IL-13 inducing peptides. GDA expands the training dataset by generating novel peptide sequences using three types of generative AI models: generative adversarial networks (GANs), diffusion models (DMs), and variational autoencoders (VAEs). The GDA framework is defined by four key parameters: the type of generative model, the sequence identity cutoff, the probability threshold (PT) for selecting generated peptides, and the augmentation ratio (AR) between generated and real peptides. Since optimizing these parameters is challenging with small datasets, we adopt a case study-oriented proof-of-concept approach using a moderately sized dataset of anti-inflammatory peptides (AIPs) to derive interpretable optimal settings. The performance of the optimized GDA was evaluated using stratified 5-fold cross-validation with cluster-based partitioning and a hold-out test on benchmark datasets. The optimized GDA was then applied to SOTA classifiers, collectively termed GDA-Pred, to identify IL-6 and IL-13 inducing peptides, both of which are limited by small dataset sizes. GDA-Pred substantially improved the prediction performance for both cytokine-inducing peptide tasks, demonstrating the feasibility of GDA-Pred as a robust and generalizable framework.

bioinformatics↗