Search bioRxiv⌕ Search

Biology subjects

Hossain, D.

Publications and source records attributed to Hossain, D..

5 recordsLinked to original sources

Benchmarking Generative Large Language Models for de novo Antibody Design and Agentic Evaluation

Despite major advances in computational antibody engineering, no systematic comparison of modern open-source LLM backbone families for antibody sequence generation exists, nor is it known whether architectural differences matter at compact model scales. In this study, five compact transformer variants inspired by prominent open-source LLM families (Llama-4, Gemma-3, DeepSeek-V3, Mistral 7B, and NVIDIA Nemotron-3) were customized and trained from scratch for de novo VH single-domain antibody (sdAb) design. All five models were pretrained from scratch on 15 million sequences from the Observed Antibody Space (OAS) database. Pretraining yielded uniformly high generative fidelity across architectures: sequence diversity 0.507-0.516 (CV=0.8%), uniqueness approaching 1.0, and novelty 0.925-0.977 (CV=2.2%). The models were subsequently fine-tuned on disease-stratified repertoires spanning SARS-CoV-2 (n=4,688), HIV (n=430), HER2 (n=22,778), and Ebola virus (n=2,868). Structural assessment of top-ranked candidates of those case studies via AlphaFold-2, Boltz-2, RoseTTAFold-2, and ESMFold produced mean pLDDT scores of 92.88{+/-}1.54 to 93.77{+/-}2.16, with no statistically significant inter-model differences (Kruskal-Wallis H=2.06, p>0.05; N=100), indicating no statistically detectable difference was observed across architectures at this compressed scale in a single-seed experiment, suggesting that generative capacity at this parameter regime is primarily determined by training data and model scale rather than family-specific design elements at this scale. Computational docking yielded predicted binding free energies of -36.34 to -65.60 kcal/mol; independent biological rigor validation through IMGT-defined CDR-H3 extraction, BLASTp novelty assessment, and NetMHCIIpan 4.3 MHC-II immunogenicity profiling collectively confirmed antigen-binding loop novelty (CDR-H3 identity 0-29% to closest database hits), germline-consistent humanness (77-90% VH germline content), and immunogenically silent antigen-binding surfaces with no strong MHC-II binders detected across CDR regions in any candidate. We further introduce a proof-of-concept agentic evaluation pipeline leveraging the Model Context Protocol (MCP) with Claude Sonnet 4.6, enabling automated structural profiling and candidate prioritization across disease targets.

bioinformatics↗

Serosurveillance Studies of Peste des Petits Ruminants (PPR) Virus in Sheep and Goats in South Asia: A Systematic Review and Meta-analysis

BackgroundPeste des Petits Ruminants (PPR), also known as the goat plague, is one of the WOAH-listed A, highly contagious and economically important viral transboundary animal diseases affecting small ruminants, and having a significant impact on the global livestock industry and international animal traffic. ObjectiveThe present study aimed to use a systematic approach to assess the pooled seroprevalence of PPRV in sheep and goats in South Asia, through a systematic review and meta-analysis of published data. MethodsA thorough search on various databases was performed to identify published research articles published between January 2000 and June 2025 reporting the seroprevalence of PPRV in small ruminants in South Asia. The articles were chosen on the basis of specific inclusion and exclusion criteria. Since the heterogeneity among the studies was significant, the pooled seroprevalence was estimated via a random effects meta-analysis model, using Stata (v19) and R software (v4.5.0). ResultsIn sheep and goats, the estimated pooled seroprevalence of PPR was 42.4% (95% CI: 35.0-49.9), whereas it was 41.8% (95% CI: 33.7-50.1) in goats and 44.5% (95% CI: 37.0- 52.0) in sheep. Subgroup analysis revealed that the pooled seroprevalence of PPRV in sheep and goat by country and vaccination status was greater in Nepal (57.0%, 95% CI: 8.4-97.7) and in vaccinated animals (57.5%, 95% CI: 47.9-66.9). ConclusionThis study highlights the need for coordinated actions, including vaccination, surveillance, and strict biosecurity, to control and eradicate the disease effectively. Moreover, authorities should adopt evidence-based strategies to support the global goal of eradicating PPR by 2030, as recommended by the WOAH.

microbiology↗

LlamaAffinity: A Predictive Antibody Antigen Binding Model Integrating Antibody Sequences with Llama3 Backbone Architecture

Antibody-facilitated immune responses are central to the bodys defense against pathogens, viruses, and other foreign invaders. The ability of antibodies to specifically bind and neutralize antigens is vital for maintaining immunity. Over the past few decades, bioengineering advancements have significantly accelerated therapeutic antibody development. These antibody-derived drugs have shown remarkable efficacy, particularly in treating Cancer, SARS-Cov-2, autoimmune disorders, and infectious diseases. Traditionally, experimental methods for affinity measurement have been time-consuming and expensive. With the realm of Artificial Intelligence, in silico medicine has revolutionized; recent developments in machine learning, particularly the use of large language models (LLMs) for representing antibodies, have opened up new avenues for AI-based designing and improving affinity prediction. Herein, we present an advanced antibody-antigen binding affinity prediction model (LlamaAffinity), leveraging an open-source Llama 3 backbone and antibody sequence data employed from the Observed Antibody Space (OAS) database. The proposed approach significantly improved over existing state-of-the-art (SOTA) approaches (AntiFormer, AntiBERTa, AntiBERTy) across multiple evaluation metrics. Specifically, the model achieved an accuracy of 0.9640, an F1-score of 0.9643, a precision of 0.9702, a recall of 0.9586, and an AUC-ROC of 0.9936. Moreover, this strategy unveiled higher computational efficiency, with a five-fold average cumulative training time of only 0.46 hours, significantly lower than previous studies. LlamaAffinity defines a new benchmark for antibody-antigen binding affinity prediction, achieving advanced performance in the immunotherapies and immunoinformatics field. Furthermore, it can effectively assess binding affinities following novel antibody design, accelerating the discovery and optimization of therapeutic candidates.

bioinformatics↗

NeSyDPP4-QSAR: A Neuro-Symbolic AI Approach for Potent DPP-4-Inhibitor Discovery in Diabetes Treatment

Diabetes Mellitus (DM) is a global epidemic and among the top ten leading causes of mortality (WHO, 2019), projected to rank seventh by 2030. The US National Diabetes Statistics Report (2021) states that 38.4 million Americans have diabetes. Dipeptidyl Peptidase-4 (DPP-4) is an FDA-approved target for type 2 diabetes mellitus (T2DM) treatment. However, current DPP-4 inhibitors are associated with adverse effects, including gastrointestinal issues, severe joint pain (FDA safety warning), nasopharyngitis, hypersensitivity, and nausea. Identifying novel inhibitors is crucial. Direct in vivo DPP-4 inhibition assessment is costly and impractical, making in silico IC50 prediction a viable alternative. Quantitative Structure-Activity Relationship (QSAR) modeling is a widely used computational approach for chemical substance assessment. We employ LTN, a neuro-symbolic approach, alongside DNN and transformers as baselines. DPP-4-related data is sourced from PubChem, ChEMBL, BindingDB, and GTP, comprising 6,563 bioactivity records (SMILES-based compounds with IC50 values) after deduplication and thresholding. A diverse set of features including descriptors (CDK Extended-PaDEL), fingerprints (Morgan), chemical language model embeddings (ChemBERTa2), LLaMa 3.2, and physicochemical properties is used to train the NeSyDPP4-QSAR model. The NeSyDPP4-QSAR model yielded the highest accuracy, incorporating CDKextended and Morgan fingerprints, with an accuracy of 0.9725, an F1-score of 0.9723, an ROC AUC of 0.9719, and an MCC of 0.9446. The performance was benchmarked against two standard baseline models: a deep neural network and a transformer. To ensure fair comparisons, DNN models used the equivalent attributes with the same dimension and network configuration as NeSyDPP4-QSAR. Our findings showed that integrating the Neuro-symbolic strategy (neural network-based learning and symbolic reasoning) holds immense potential for discovering drugs that can inhibit diabetes mellitus and classifying biological activities that inhibit it.

bioinformatics↗

hERG-LTN: A New Paradigm in hERG Cardiotoxicity Assessment Using Neuro-Symbolic and Generative AI Embedding (MegaMolBART, Llama3.2, Gemini, DeepSeek) Approach

Assessing adverse drug reactions (ADRs) during drug development is essential for ensuring the safety of new compounds. The blockade of the Ether-a-go-go-related gene (hERG) channel plays a critical role in cardiac repolarization. Computational predictions of hERG inhibition can help foresee drug safety, but current data-driven approaches have limitations. Therefore, a new paradigm that bridges the gap between data and knowledge offers an alternative for advancing precision pharmacogenomics in assessing hERG cardiotoxicity. This study aims to develop a reasoning-based, in silico, robust model for predicting drug-induced hERG inhibition, facilitating new drug development by reducing time and cost, supporting downstream in vitro and in vivo testing. In this study, we constructed a new cohort, UnihERG_DB, by sourcing data from ChEMBL, PubChem, BindingDB, GTP, hERG Karims, and hERG Blockers bioactivity databases. The final dataset comprises 20,409 structures represented as SMILES (Simplified Molecular Input Line Entry System), labeled as hERG blockers (IC50 < 10 {micro}M) or non-hERG blockers (IC50 [&ge;] 10 {micro}M). Molecular features were extracted using Morgan and CDK fingerprints. Furthermore, we explored embedding feature computation using cutting-edge Large Language Models, including NVIDIA MegaMolBART, LLaMA 3.2, Gemini, and DeepSeek. Finally, we utilized the Logic Tensor Network (LTN), an advanced AI framework, to train and develop the hERG predictive model. Model performance was evaluated using two benchmarks: External Test-1 and hERG-70. The Logic Tensor Network (LTN) outperformed several models, including CardioTox, M-PNN, DeepHIT, CardPred, OCHEM Predictor-II, Pred-hERG 4.2, Random Forest, and Gradient Boosting. On the External Test-1 dataset, LTN achieved an accuracy of 0.931, a specificity of 0.928, and a sensitivity of 0.933. Furthermore, on the hERG-70 benchmark, LTN achieved an accuracy (ACC) of 0.827, a specificity (SPE) of 0.890, and a correct classification rate (CCR) of 0.833. Overall, the Neuro-Symbolic AI approach sets a new standard for hERG-related cardiotoxicity assessment, yielding competitive results with current state-of-the-art (SOTA) models, and highlights its potential for advancing precision pharmacogenomics in drug discovery and development (GitHub).

bioinformatics↗