Search bioRxivSearch

Biology subjects

Guigo, R.

Publications and source records attributed to Guigo, R..

6 recordsLinked to original sources

Nearly all new protein-coding predictions in the CHESS database are not protein-coding

In a 2018 paper posted to bioRxiv, Pertea et al. presented the CHESS database, a new catalog of human gene annotations that includes 1,178 new protein-coding predictions. These are based on evidence of transcription in human tissues and homology to earlier annotations in human and other mammals. Here, we reanalyze the evidence used by CHESS, and find that nearly all protein-coding predictions are false positives. We find that 86% overlap transposons marked by RepeatMasker that are known to frequently result in false positive protein-coding predictions. More than half are homologous to only nine Alu-derived primate sequences corresponding to an erroneous and previously withdrawn Pfam protein domain. The entire set shows poor evolutionary conservation and PhyloCSF protein-coding evolutionary signatures indistinguishable from noncoding RNAs, indicating lack of protein-coding constraint. Only four predictions are supported by mass spectrometry evidence, and even those matches are inconclusive. Overall, the new protein-coding predictions are unsupported by any credible experimental or evolutionary evidence of function, result primarily from homology to genes incorrectly classified as protein-coding, and are unlikely to encode functional proteins.

genomics

A systematic survey of human tissue-specific gene expression and splicing reveals new opportunities for therapeutic target identification and evaluation

Differences in the expression of genes and their splice isoforms across human tissues are fundamental factors to consider for therapeutic target evaluation. To this end, we conducted a transcriptome-wide survey of tissue-specific gene expression and splicing events in the unprecedented collection of 8527 high-quality RNA-seq samples from the GTEx project, covering 36 human peripheral tissues and 13 brain subregions. We derived a weighted tissue-specificity scoring scheme accounting for the similarity of related tissues and inherent variability across individual samples. We showed that ~50.6% of all annotated human genes show tissue-specific expression, including many low abundance transcripts vastly underestimated by previous array-based expression atlases. As utilities for drug discovery, we demonstrated that tissue-specificity is a highly desirable attribute of validated drug targets and tissue-specificity can be used to prioritize disease-associated genes from genome-wide association studies (GWAS). Using brain striatum-specific gene expression as an example, we provided a template to leverage tissue-specific gene expression to identify novel therapeutic targets. Mining of tissue-specific splicing further reveals new opportunities for tissue-specific targeting. Thus, the high quality transcriptome atlas provided by the GTEx is an invaluable resource for drug discovery and systematic analysis anchored on the human tissue specific gene expression provides a promising avenue to identify novel therapeutic target hypotheses.

genomics

Modified penetrance of coding variants by cis-regulatory variation shapes human traits

Coding variants represent many of the strongest associations between genotype and phenotype, however they exhibit inter-individual differences in effect, known as variable penetrance. In this work, we study how cis-regulatory variation modifies the penetrance of coding variants in their target gene. Using functional genomic and genetic data from GTEx, we observed that in the general population, purifying selection has depleted haplotype combinations that lead to higher penetrance of pathogenic coding variants. Conversely, in cancer and autism patients, we observed an enrichment of haplotype combinations that lead to higher penetrance of pathogenic coding variants in disease implicated genes, which provides direct evidence that regulatory haplotype configuration of causal coding variants affects disease risk. Finally, we experimentally demonstrated that a regulatory variant can modify the penetrance of a coding variant by introducing a Mendelian SNP using CRISPR/Cas9 on distinct expression haplotypes and using the transcriptome as a phenotypic readout. Our results demonstrate that joint effects of regulatory and coding variants are an important part of the genetic architecture of human traits, and contribute to modified penetrance of disease-causing variants.

genetics

Unique genomic features and deeply-conserved functions of long non-coding RNAs in the Cancer LncRNA Census (CLC)

Long non-coding RNAs (lncRNAs) that drive tumorigenesis are a growing focus of cancer genomics studies. To facilitate further discovery, we have created the \"Cancer LncRNA Census\" (CLC), a manually-curated and strictly-defined compilation of lncRNAs with causative roles in cancer. CLC has two principle applications: first, as a resource for training and benchmarking de novo identification methods; and second, as a dataset for studying the fundamental properties of these genes.\n\nCLC Version 1 comprises 122 lncRNAs implicated in 29 distinct cancers. LncRNAs are included based on functional or genetic evidence for causative roles in cancer progression. All belong to the GENCODE reference annotation, to enable integration across projects and datasets. For each entry, the evidence type, biological activity (oncogene or tumour suppressor), source reference and cancer type are recorded. Supporting its usefulness, CLC genes are significantly enriched amongst de novo predicted driver genes from PCAWG. CLC genes are distinguished from other lncRNAs by a series of features consistent with biological function, including gene length, high expression and sequence conservation of both exons and promoters. We identify a trend for CLC genes to be co-localised with known protein-coding cancer genes along the human genome. Finally, by integrating data from transposon-mutagenesis functional screens, we show that mouse orthologues of CLC genes tend also to be cancer genes.\n\nThus CLC represents a valuable resource for research into long non-coding RNAs in cancer. Their evolutionary and genomic properties have implications for understanding disease mechanisms and point to conserved functions across ~80 million years of evolution.

bioinformatics

LncATLAS database for subcellular localisation of long noncoding RNAs

BackgroundThe subcellular localisation of long noncoding RNAs (lncRNAs) holds valuable clues to their molecular function. However, measuring localisation of newly-discovered lncRNAs involves time-consuming and costly experimental methods.\n\nResultsWe have created \"LncATLAS\", a comprehensive resource of lncRNA localisation in human cells based on RNA-sequencing datasets. Altogether, 6768 GENCODE-annotated lncRNAs are represented across various compartments of 15 cell lines. We introduce \"Relative concentration index\" (RCI) as a useful measure of localisation derived from ensemble RNAseq measurements. LncATLAS is accessible through an intuitive and informative webserver, from which lncRNAs of interest are accessed using identifiers or names. Localisation is presented across cell types and organelles, and may be compared to the distribution of all other genes. Publication-quality figures and raw data tables are automatically generated with each query, and the entire dataset is also available to download.\n\nConclusionsLncATLAS makes lncRNA subcellular localisation data available to the widest possible number of researchers. It is available at lncATLAS.crg.eu.

bioinformatics

High-throughput annotation of full-length long noncoding RNAs with Capture Long-Read Sequencing (CLS)

Accurate annotations of genes and their transcripts is a foundation of genomics, but no annotation technique presently combines throughput and accuracy. As a result, reference gene collections remain incomplete: many gene models are fragmentary, while thousands more remain uncatalogued-particularly for long noncoding RNAs (lncRNAs). To accelerate lncRNA annotation, the GENCODE consortium has developed RNA Capture Long Seq (CLS), combining targeted RNA capture with third-generation long-read sequencing. We present an experimental re-annotation of the GENCODE intergenic lncRNA population in matched human and mouse tissues, resulting in novel transcript models for 3574 / 561 gene loci, respectively. CLS approximately doubles the annotated complexity of targeted loci, outperforming existing short-read techniques. Full-length transcript models produced by CLS enable us to definitively characterize the genomic features of lncRNAs, including promoter- and gene-structure, and protein-coding potential. Thus CLS removes a longstanding bottleneck of transcriptome annotation, generating manual-quality full-length transcript models at high-throughput scales.\n\nAbbreviations

genomics