Search bioRxiv⌕ Search

Biology subjects

Bhattacharya, H.

Publications and source records attributed to Bhattacharya, H..

2 recordsLinked to original sources

A structural antibody benchmark for leakage-aware deep-learning evaluation

Deep-learning methods for antibody structure prediction, antibody-antigen interaction modelling and design are advancing rapidly. However, comparisons across studies remain difficult because training and test sets are often constructed independently, and a temporal cutoff alone does not prevent train-test leakage. We present SABLE (Structural Antibody Benchmark for deep-Learning Evaluation), a versioned structural antibody resource that couples a fixed training collection with a leakage-controlled held-out test set for reproducible machine-learning development and evaluation. SABLE combines 16,511 experimental training entries with 327 manually reviewed test entries and 3,274 high-confidence, patent-derived AlphaFold3 models spanning 476 antigens. Candidate test structures were selected after the AlphaFold3 temporal cutoff and filtered against the training set using antibody and antigen sequence similarity filters. Each test entry records its nearest training set neighbour, allowing users to quantify remaining relatedness and stratify performance by similarity. Versioned releases provide metadata, processed structures, CDR annotations, redundancy labels and model-confidence fields. A Python/PyTorch API and standardised benchmark metrics provide reproducible database access and evaluation code without requiring additional antibody-structure processing.

bioinformatics↗

PHLAME: A benchmark for continuous evaluation of host phenotype prediction from shotgun metagenomic data

Predicting host phenotypes from shotgun metagenomic data is essential for translating microbiome research into clinical practice)Despite the development of numerous computational tools for this task, researchers often default to traditional machine learning methods such as Random Forest)This hesitancy to adopt newer methods stems from their complexity as well as the lack of standardized evaluations, as most tools are assessed on different datasets and compared against a limited set of methods)Here, we introduce LAMPP, a standardized benchmark for evaluating methods for predicting host phenotypes from gut metagenomic data)LAMPP features a diverse range of prediction tasks and enables consistent, comparative assessments across prediction tools)Our systematic evaluation of existing tools shows that classic machine learning methods (e.g., Random Forest) perform competitively, offering both ease of use and state-of-the-art results)At the same time, it demonstrates that microbiome-based phenotype prediction remains a challenging problem)By providing a consistent platform for ongoing evaluation and access to raw sequencing data, LAMPP motivates the development of novel prediction pipelines from raw sequencing data to phenotype prediction, including novel sample representation and data augmentation strategies)LAMPP is publicly available for ongoing benchmarking at https://lampp.yassourlab.com/.

bioinformatics↗