Search bioRxiv⌕ Search

Biology subjects

Cerdan-Velez, D.

Publications and source records attributed to Cerdan-Velez, D..

3 recordsLinked to original sources

More than 2,500 coding genes in the human reference gene set still have unsettled status

In 2018 we analysed the three main repositories for the human proteome, Ensembl/GENCODE, RefSeq and UniProtKB. They disagreed on the coding status of one of every eight annotated coding genes. The analysis inspired bilateral collaborations between annotation groups. Here we have repeated our analysis with updated versions of the three reference coding gene sets. Superficially, little appears to have changed. Although there are slightly fewer genes predicted as coding overall, the three groups still disagree on the status of 2,606 annotated genes. However, a comparison without read-through genes and immunoglobulin fragments shows that the three reference sets have merged or reclassified more than 700 genes since the last analysis and that just 0.6% of Ensembl/GENCODE coding genes are not also annotated by the other two reference sets. We used eight features indicative of non-coding genes to examine the 21,873 coding genes annotated across the three reference sets. We found that more than 2,000 had one or more potential non-coding features. While some of these genes will be protein coding, we believe that most are likely to be non-coding genes or pseudogenes. Our results suggest that annotators still vastly overestimate the number of true coding genes.

genomics↗

A deep audit of the PeptideAtlas database uncovers evidence for unannotated coding genes and aberrant translation

The human genome has been the subject of intense scrutiny by experimental and manual curation projects for more than two decades. Novel coding genes have been proposed from large-scale RNASeq, ribosome profiling and proteomics experiments. Here we carry out an in-depth analysis of an entire proteomics database. We analysed the proteins, peptides and spectra housed in the human build of the PeptideAtlas proteomics database to identify coding regions that are not yet annotated in the GENCODE reference gene set. We find support for hundreds of missing alternative protein isoforms and unannotated upstream translations, and evidence of cross-contamination from other species. There was reliable peptide evidence for 34 novel unannotated open reading frames (ORFs) in PeptideAtlas. We find that almost half belong to coding genes that are missing from GENCODE and other reference sets. Most of the remaining ORFs were not conserved beyond human, however, and their peptide confirmation was restricted to cancer cell lines. We show that this is strong evidence for aberrant translation, raising important questions about the extent of aberrant translation and how these ORFs should be annotated in reference genomes.

genomics↗

GENCODE: massively expanding the lncRNA catalog through capture long-read RNA sequencing

Accurate and complete gene annotations are indispensable for understanding how genome sequences encode biological functions. For more than twenty years, the GENCODE consortium has developed reference annotations for the human and mouse genomes, becoming a foundation for biomedical and genomics communities worldwide. Nevertheless, collections of important yet poorly-understood gene classes like long non-coding RNAs (lncRNAs) remain incomplete and scattered across multiple, uncoordinated catalogs. To address this, GENCODE has undertaken the most comprehensive lncRNA annotation effort to date. This is founded on the manually supervised computational annotation of full-length targeted long-read sequencing, on matched embryonic and adult tissues, of orthologous regions in human and mouse. Altogether 17,931 human genes (140,268 transcripts) and 22,784 mouse genes (136,169 transcripts) have been added to the GENCODE catalog representing a 2-fold and 6-fold growth in transcripts, respectively - the greatest increase in the number of annotated human genes since the sequencing of the human genome. Our targeted design assigned human-mouse orthologs at a rate beyond previous studies, tripling the number of human disease-associated lncRNAs that have mouse orthologs. Novel lncRNA genes consistently exhibit biological signals of functionality, and they greatly enhance the functional interpretability of the human genome. While poorly expressed in bulk RNA-Seq samples, many of them are highly expressed in specific cell populations, maybe even contributing to cell-type determination. The expanded GENCODE lncRNA annotations mark a critical step toward deciphering the human and mouse genomes.

genomics↗