Search bioRxiv⌕ Search

Biology subjects

Vasquez-Rios, C.

Publications and source records attributed to Vasquez-Rios, C..

2 recordsLinked to original sources

Utanos: A general-purpose shallow whole-genome sequencing analysis workflow identifies interpretable copy number signatures

SummaryA modular FASTQ-to-figures solution for analyzing low-depth or shallow whole-genome sequencing (sWGS) data. Shallow WGS can be used to detect copy number (CN) aberrations, Homologous Recombination Deficiency (HRD), and to detect and create CN signatures. Growing in popularity, this sequencing type is used for neonatal diagnostics and studying cancer. One of the major benefits is the reduced cost compared with deeper sequencing modalities such as Whole Genome Sequencing (WGS). With just 15 million reads often targeted, and sample substrate options like Formalin-Fixed, Paraffin-Embedded (FFPE) blocks widely available, this approach enables an affordable study to be performed at scale. Our pipeline and R package are an end-to-end solution implemented with reusability and modularity in mind. It makes entry and exit from the ecosystem easy, providing regular standardized output formats throughout execution. The pipeline is written in the well-supported and cross-platform Nextflow framework and has been submitted for inclusion in nf-core. Additionally, a Docker image for the utanos R package has been created to improve modularity. Availability and ImplementationThe latest version of all software is freely available on GitHub. For the full processing pipeline, visit: https://github.com/Huntsmanlab/swgs-processing-pipeline. For just the utanos R package, visit: https://github.com/Huntsmanlab/utanos.

bioinformatics↗

Generating cognate epitope sequences of T-cell receptors with a generative transformer

Single-cell TCR sequencing enables high-resolution analysis of T Cell Receptor (TCR) diversity and clonality, offering valuable insights into immune responses and disease mechanisms. However, identifying cognate epitopes for individual TCRs requires complex and costly functional assays. We address this challenge with EpitopeGen, a largescale transformer model based on the GPT-2 architecture that generates potential cognate epitope sequences directly from TCR sequences. To overcome the scarcity of TCR-epitope binding pairs ({approx} 100, 000), EpitopeGen uses a semi-supervised learning method, termed BINDSEARCH, which searches over 70 billion potential pairs and incorporates high binding affinity predictions as pseudo-labels. To incorporate CD8+ T cell biology into the model as an inductive bias, EpitopeGen employs a novel data balancing method, termed Antigen Category Filter, that carefully controls antigen category ratios in its training dataset. EpitopeGen significantly outperforms baseline approaches, generating epitopes with high binding affinity, diversity, naturalness, and biophysical stability. Notably, the epitopes generated by EpitopeGen follow biologically plausible antigen category distributions, a crucial feature not achieved by other models. Using EpitopeGen, we directly identify subsets of clonally expanded tumorinfiltrating lymphocytes that recognize tumor-associated antigens, exhibiting elevated cytotoxicity and reduced exhaustion markers. From COVID-19 patients, EpitopeGen detects T cells that recognize COVID-19 spike proteins and non-structural proteins with distinct transcriptomic characteristics. In conclusion, EpitopeGen represents the first computational method that enables direct inference of antigen recognition profiles of CD8+ T cells from plain TCR repertoires.

immunology↗