Search bioRxiv⌕ Search

Biology subjects

Vara, C.

Publications and source records attributed to Vara, C..

2 recordsLinked to original sources

Identification of a large class of cancer-germline microproteins as a source of immunotherapeutic targets

Classical cancer germline-antigens (CGAs) are proteins that are expressed in the male germinal line but not in somatic tissues, and that can also become expressed in tumors. However, the vast majority of testis-specific transcripts are long non-coding RNAs (lncRNAs) rather than protein-coding genes. Since recent studies have shown that many lncRNAs contain non-canonical open reading frames (ncORFs) that are translated into small proteins, or microproteins, there could be a large class of non-canonical cancer-germline antigens (ncCGAs) that remains to be discovered. Here, we integrate ribosome profiling from human testis and cancer cell lines with paired tumor/normal transcriptomes from 917 patients across eight common cancer types to define a comprehensive catalog of ncCGAs. This set comprises 235 ncCGAs encoded by lncRNAs or mRNA untranslated regions (5UTRs and 3UTRs), compared to 192 canonical CGAs (cCGAs) with similar expression patterns. We show that ncCGAs are evolutionary young, consistent with recent de novo emergence in the rapidly evolving male germline. Moreover, a large fraction is expressed across multiple patients and cancer types, indicating recurrent reactivation mechanisms in tumors. We further find that ncCGAs are frequently located in cancer-amplified regions or associated with MYC or E2F-regulated pathways, which may explain their expression in cancer. Finally, we provide strong evidence that a subset of ncCGAs give rise to potentially immunogenic HLA class I bound peptides. Together, our results describe a previously unexplored class of tumor-restricted antigens with potential applications in cancer immunotherapy.

cancer biology↗

GENCODE: massively expanding the lncRNA catalog through capture long-read RNA sequencing

Accurate and complete gene annotations are indispensable for understanding how genome sequences encode biological functions. For more than twenty years, the GENCODE consortium has developed reference annotations for the human and mouse genomes, becoming a foundation for biomedical and genomics communities worldwide. Nevertheless, collections of important yet poorly-understood gene classes like long non-coding RNAs (lncRNAs) remain incomplete and scattered across multiple, uncoordinated catalogs. To address this, GENCODE has undertaken the most comprehensive lncRNA annotation effort to date. This is founded on the manually supervised computational annotation of full-length targeted long-read sequencing, on matched embryonic and adult tissues, of orthologous regions in human and mouse. Altogether 17,931 human genes (140,268 transcripts) and 22,784 mouse genes (136,169 transcripts) have been added to the GENCODE catalog representing a 2-fold and 6-fold growth in transcripts, respectively - the greatest increase in the number of annotated human genes since the sequencing of the human genome. Our targeted design assigned human-mouse orthologs at a rate beyond previous studies, tripling the number of human disease-associated lncRNAs that have mouse orthologs. Novel lncRNA genes consistently exhibit biological signals of functionality, and they greatly enhance the functional interpretability of the human genome. While poorly expressed in bulk RNA-Seq samples, many of them are highly expressed in specific cell populations, maybe even contributing to cell-type determination. The expanded GENCODE lncRNA annotations mark a critical step toward deciphering the human and mouse genomes.

genomics↗