Search bioRxiv⌕ Search

Biology subjects

Cohen-Lavi, L.

Publications and source records attributed to Cohen-Lavi, L..

4 recordsLinked to original sources

BenchRep-T: A Systematic Evaluation of T-Cell Repertoire-Based Disease Diagnostics

Adaptive immune receptor repertoire sequencing data has emerged as a promising potential modality for disease diagnosis, relying on computational methods to analyze T-cell receptor (TCR) sequences from an individuals blood sample. Published methods rely on different cohorts, data preprocessing pipelines, and evaluation metrics, making direct comparison across methods challenging. We present BenchRep-T, a unified benchmark that standardizes multiple publicly available TCR repertoire datasets and evaluates nine computational approaches, spanning statistical enrichment of shared sequences, feature-engineered ensembles, deep learning, and sequence clustering. BenchRep-T evaluates methods on four tasks: disease classification across conditions, performance scaling under restricted sequence-sampling depth, recovery of known antigen-specific driver sequences, and evaluation of sensitivity to demographic confounding. Under controlled evaluation, simple baselines prove competitive, with tree-based models trained on V- and J-gene usage and short sequence motifs approaching the classification performance of more complex methods. Our findings underscore the complexity of modeling TCR repertoire data, and show that no single method dominates across all tasks. BenchRep-T provides a framework for rigorous and reproducible evaluation of TCR repertoire classification methods to accelerate the development of immune repertoire-based diagnostics.

immunology↗

Computational design of HLA class I superbinders for broad T cell immunogenicity

Human leukocyte antigen (HLA) class I molecules are highly polymorphic, restricting peptide binding to narrow sequence subsets. Designing peptides that bind multiple HLA supertypes-- termed superbinders--offers a promising strategy for broad-spectrum T-cell vaccines and immunotherapies. Here we present superHLA, a computational framework that combines Markov Chain Monte Carlo optimization with state-of-the-art MHC binding predictors to design synthetic 9-mer peptides with broad HLA-binding profiles. Using superHLA, we generated over 190,000 candidate superbinders predicted to bind 8-12 HLA class I alleles across distinct supertypes. A multi-tier filtering pipeline--incorporating sequence clustering, synthesis feasibility, cross-predictor validation, and self-peptidome exclusion--yielded a final panel of 100 peptides for experimental testing. Of these, 21 bound [≥]4 supertypes in vitro, including one that bound 9. Superbinders displayed distinct anchor residue preferences and showed minimal similarity to human peptides. These results suggest that HLA superbinders are more abundant than previously recognized and can be rationally designed at scale. This approach supports development of pan-HLA immunogens with broad population coverage and potential applications in vaccine design, neoantigen discovery, and immunotherapy.

immunology↗

Benchmarking solutions to the T-cell receptor epitope prediction problem: IMMREP22 workshop report

Many different solutions to predicting the cognate epitope target of a T-cell receptor (TCR) have been proposed. However several questions on the advantages and disadvantages of these different approaches remain unresolved, as most methods have only been evaluated within the context of their initial publications and data sets. Here, we report the findings of the first public TCR-epitope prediction benchmark performed on 23 prediction models in the context of the ImmRep 2022 TCR-epitope specificity workshop. This benchmark revealed that the use of paired-chain alpha-beta, as well as CDR1/2 or V/J information, when available, improves classification obtained with CDR3 data, independent of the underlying approach. In addition, we found that straight-forward distance-based approaches can achieve a respectable performance when compared to more complex machine-learning models. Finally, we highlight the need for a truly independent follow-up benchmark and provide recommendations for the design of such a next benchmark.

bioinformatics↗

TCR meta-clonotypes for biomarker discovery with tcrdist3: quantification of public, HLA-restricted TCR biomarkers of SARS-CoV-2 infection

As the mechanistic basis of adaptive cellular antigen recognition, T cell receptors (TCRs) encode clinically valuable information that reflects prior antigen exposure and potential future response. However, despite advances in deep repertoire sequencing, enormous TCR diversity complicates the use of TCR clonotypes as clinical biomarkers. We propose a new framework that leverages antigen-enriched repertoires to form meta-clonotypes - groups of biochemically similar TCRs - that can be used to robustly identify and quantify functionally similar TCRs in bulk repertoires. We apply the framework to TCR data from COVID-19 patients, generating 1831 public TCR meta-clonotypes from the 17 SARS-CoV-2 antigen-enriched repertoires with the strongest evidence of HLA-restriction. Applied to independent cohorts, meta-clonotypes targeting these specific epitopes were more frequently detected in bulk repertoires compared to exact amino acid matches, and 59.7% (1093/1831) were more abundant among COVID-19 patients that expressed the putative restricting HLA allele (FDR < 0.01), demonstrating the potential utility of meta-clonotypes as antigen-specific features for biomarker development. To enable further applications, we developed an open-source software package, tcrdist3, that implements this framework and facilitates flexible workflows for distance-based TCR repertoire analysis.

immunology↗