Search bioRxiv⌕ Search

Biology subjects

Anwar, A. M.

Publications and source records attributed to Anwar, A. M..

3 recordsLinked to original sources

A unified framework for batch correction and missing data handling in large-scale and single-cell mass spectrometry proteomics

Large-scale mass spectrometry (MS)-based proteomics, including single-cell proteomics, is routinely affected by technical variation arising from discrete batch effects, inter-laboratory differences and continuous signal drift during data acquisition. Current correction strategies typically address these sources of unwanted variation independently and often require either removal of proteins with missing values or imputation before correction, both of which may lead to information loss and potential amplification of technical bias. Here we present NMFBatch, a unified statistical framework that simultaneously models discrete and continuous unwanted variation in bulk and single-cell proteomics data. NMFBatch integrates non-negative matrix factorization with generalized additive modelling and directly accommodates missing values, thereby enabling both on-the-fly imputation during correction and optional post-correction imputation. Benchmarking against six batch-correction methods using multi-laboratory reference datasets and a large plasma proteomics cohort, shows that NMFBatch consistently reduces batch-associated variation while preserving biological structure under both balanced and confounded experimental designs. Application to single-cell proteomics data further showed effective reduction of TMT- and acquisition-associated variation while retaining biologically meaningful clustering. Together, these results establish NMFBatch as a flexible framework for modelling unwanted variation in proteomics experiments, with potential applications in cross-cohort harmonization and integrative proteomics analysis. Graphical AbstractCreated in BioRender. Youssef, A. (2026) https://BioRender.com/c1q1yxt O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=181 SRC="FIGDIR/small/726178v2_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@2b7cd1org.highwire.dtl.DTLVardef@10fada3org.highwire.dtl.DTLVardef@50e66corg.highwire.dtl.DTLVardef@147f81c_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

LimROTS: A Hybrid Method Integrating Empirical Bayes and Reproducibility-Optimized Statistics for Robust Analysis of Proteomics Data

MotivationDifferential expression analysis plays a vital role in omics research enabling precise identification of features that associate with different phenotypes. This process is critical for uncovering biological differences between conditions, such as disease versus healthy states. In proteomics, several statistical methods have been used, ranging from simple t-tests to more advanced methods like limma and ROTS. However, a flexible method for reproducibility-optimized statistics tailored for clinical omics data has been lacking. ResultsIn this study, we developed LimROTS, a hybrid method integrating the linear model and empirical Bayes method from the limma framework with the Reproducibility-Optimized Statistics from ROTS, to create a novel moderated ranking statistic, for robust and flexible analysis of proteomics data. We validated its performance using twenty-one proteomics gold standard spike-in datasets with different protein mixtures, MS instruments, and techniques for benchmarking. This hybrid approach improves accuracy and reproducibility of complex proteomics data, making LimROTS a powerful tool for high-dimensional omics data analysis. Availability and ImplementationLimROTS has been implemented as an R/Bioconductor package, available at https://bioconductor.org/packages/LimROTS/. Additionally, the code used in this study is available in GitHub repository https://github.com/AliYoussef96/LimROTSmanuscript Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=181 SRC="FIGDIR/small/636801v2_ufig1.gif" ALT="Figure 1"> View larger version (47K): org.highwire.dtl.DTLVardef@1e6dd66org.highwire.dtl.DTLVardef@1d1589eorg.highwire.dtl.DTLVardef@110f73corg.highwire.dtl.DTLVardef@d78ef2_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Insights into The Codon Usage Bias of 13 Severe Acute Respiratory Syndrome Coronavirus 2 (SARS-CoV-2) Isolates from Different Geo-locations

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is the causative agent of Coronavirus disease 2019 (COVID-19) which is an infectious disease that spread throughout the world and was declared as a pandemic by the World Health Organization (WHO). In this study, we performed a genome-wide analysis on the codon usage bias (CUB) of 13 SARS-CoV-2 isolates from different geo-locations (countries) in an attempt to characterize it, unravel the main force shaping its pattern, and understand its adaptation to Homo sapiens. Overall results revealed that, SARS-CoV-2 codon usage is slightly biased similarly to other RNA viruses. Nucleotide and dinucleotide compositions displayed a bias toward A/U content in all codon positions and CpU-ended codons preference, respectively. Eight common putative preferred codons were identified, and all of them were A/U-ended (U-ended: 7, A-ended: 1). In addition, natural selection was found to be the main force structuring the codon usage pattern of SARS-CoV-2. However, mutation pressure and other factors such as compositional constraints and hydrophobicity had an undeniable contribution. Two adaptation indices were utilized and indicated that SARS-CoV-2 is moderately adapted to Homo sapiens compared to other human viruses. The outcome of this study may help in understanding the underlying factors involved in the evolution of SARS-CoV-2 and may aid in vaccine design strategies.

genomics↗