Search bioRxiv⌕ Search

Biology subjects

Chapin, S. R.

Publications and source records attributed to Chapin, S. R..

3 recordsLinked to original sources

Sparse, random sampling is sufficient for central tolerance

Negative selection in the thymus limits autoimmunity by eliminating T cells that react strongly to self. Individual T cells, however, are only exposed to a small fraction of all self peptides during their "training" in the thymus, and it is puzzling how tolerance can be generalized to the remaining "test" self peptides across peripheral tissues in the body. Using a machine learning perspective, we show that such generalization is possible because the immune system satisfies two conditions: first that peptide abundance levels in the human thymus and periphery are highly correlated (i.e., training distribution {approx} test distribution), and second that cross-reactivity allows T cells to effectively learn binding information of similar peptides without explicitly interacting with all of them. Together, we show that sparse, random sampling of only 10% of self peptides in the thymus is sufficient to avoid reactivity to 90% of peripheral self, and we support this result with diverse experimental data. We then validate two predictions by our model; the first is that only 200-250 antigen presenting cells need to be seen by a T cell to ensure its robust selection, and the second relates how peptides missing from the thymus can drive auto-immunity of peripheral tissues. Overall, we provide a plausible answer to a long-standing question underlying adaptive immunity, and we highlight how generalization, a fundamental challenge faced by nearly every learning algorithm, is uniquely tackled by the immune system.

immunology↗

BATMAN: Improved T cell receptor cross-reactivity prediction benchmarked on a comprehensive mutational scan database

Predicting T cell receptor (TCR) activation is challenging due to the lack of both unbiased benchmarking datasets and computational methods that are sensitive to small mutations to a peptide. To address these challenges, we curated a comprehensive database, called BATCAVE, encompassing complete single amino acid mutational assays of more than 22,000 TCR-peptide pairs, centered around 25 immunogenic human and mouse epitopes, across both major histocompatibility complex classes, against 151 TCRs. We then present an interpretable Bayesian model, called BATMAN, that can predict the set of peptides that activates a TCR. We also developed an active learning version of BATMAN, which can efficiently learn the binding profile of a novel TCR by selecting an informative yet small number of peptides to assay. When validated on our database, BATMAN outperforms existing methods and reveals important biochemical predictors of TCR-peptide interactions. Finally, we demonstrate the broad applicability of BATMAN, including for predicting off-target effects for TCR-based therapies and polyclonal T cell responses.

immunology↗

copepodTCR: Identification of Antigen-Specific T Cell Receptors with combinatorial peptide pooling

T cell receptor (TCR) repertoire diversity enables the antigen-specific immune responses against the vast space of possible pathogens. Identifying TCR-antigen binding pairs from the large TCR repertoire and antigen space is crucial for biomedical research. Here, we introduce copepodTCR, an open-access tool to design and interpret high-throughput experimental TCR specificity assays. copepodTCR implements a combinatorial peptide pooling scheme for efficient experimental testing of T cell responses against large overlapping peptide libraries, that can be used to identify the specificity of (or "deorphanize") TCRs. The scheme detects experimental errors and, coupled with a hierarchical Bayesian model for unbiased interpretation, identifies the response-eliciting peptide sequence for a TCR of interest out of hundreds of peptides tested using a simple experimental set-up. Using in silico simulations, we demonstrate the varied experimental settings in which copepodTCR yields efficient and interpretable TCR specificity results. We validated our approach on a library of 253 overlapping peptides covering the SARS-CoV-2 spike protein, split across 12 pools. A single stimulation with combinatorial pools identified the correct epitope of two TCRs with known specificity and then deorphanized two SARS-CoV-2 associated TCRs shared among a large cohort of COVID-19 patients. We provide experimental guides to efficiently design larger screens covering thousands of peptides which will be crucial to identify antigen-specific T cells and their targets from limited clinical material.

immunology↗