Search bioRxivSearch

bioRxiv · 10.1101/128728

Leveraging The Resolution Of RNA-Seq Markedly Increases The Number Of Causal eQTLs And Candidate Genes In Human Autoimmune Disease

Abstract

Genome-wide association studies have identified hundreds of risk loci for autoimmune disease, yet only a minority ([~]25%) share genetic effects with changes to gene expression (eQTLs) in immune cells. RNA-Seq based quantification at whole-gene resolution, where abundance is estimated by culminating expression of all transcripts or exons of the same gene, is likely to account for this observed lack of colocalisation as subtle isoform switches and expression variation in independent exons can be concealed. We performed integrative cis-eQTL analysis using association statistics from twenty autoimmune diseases (560 independent loci) and RNA-Seq data from 373 individuals of the Geuvadis cohort profiled at gene-, isoform-, exon-, junction-, and intron-level resolution in lymphoblastoid cell lines. After stringently testing for a shared causal variant using both the Joint Likelihood Mapping and Regulatory Trait Concordance frameworks, we found that gene-level quantification significantly underestimated the number of causal cis-eQTLs. Only 5.0-5.3% of loci were found to share a causal cis-eQTL at gene-level compared to 12.9-18.4% at exon-level and 9.6-10.5% at junction-level. More than a fifth of autoimmune loci shared an underlying causal variant in a single cell type by combining all five quantification types; a marked increase over current estimates of steady-state causal cis-eQTLs. As an example, we dissected in detail the genetic associations of systemic lupus erythematosus and functionally annotated the candidate genes. Many of the known and novel genes were concealed at gene-level (e.g. BANK1, UBE2L3, IKZF2, TYK2, LYST). By leveraging RNA-Seq, we were able to isolate the specific transcripts, exons, junctions, and introns modulated by the cis-eQTL - which supports the targeted design of follow-up functional studies involving alternative splicing. Causal cis-eQTLs detected at different quantification types were also found to localise to discrete epigenetic annotations. We provide our findings from all twenty autoimmune diseases as a web resource.\n\nAuthor SummaryIt is well acknowledged that non-coding genetic variants contribute to disease susceptibility through alteration of gene expression levels (known as eQTLs). Identifying the variants that are causal to both disease risk and changes to expression levels has not been easy and we believe this is in part due to how expression is quantified using RNA-Sequencing (RNA-Seq). Whole-gene expression, where abundance is estimated by culminating expression of all transcripts or exons of the same gene, is conventionally used in eQTL analysis. This low resolution may conceal subtle isoform switches and expression variation in independent exons. Using isoform-, exon-, and junction-level quantification can not only point to the candidate genes involved, but also the specific transcripts implicated. We make use of existing RNA-Seq expression data profiled at gene-, isoform-, exon-, junction-, and intron-level, and perform eQTL analysis using association data from twenty autoimmune diseases. We find exon-, and junction-level thoroughly outperform gene-level analysis, and by leveraging all five quantification types, we find >20% of autoimmune loci share a single genetic effect with gene expression. We highlight that existing and new eQTL cohorts using RNA-Seq should profile expression at multiple resolutions to maximise the ability to detect causal eQTLs and candidate genes.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Odhams, C. A., Cunninghame Graham, D. S., Vyse, T. J.. 2017-04-19. Leveraging The Resolution Of RNA-Seq Markedly Increases The Number Of Causal eQTLs And Candidate Genes In Human Autoimmune Disease. https://doi.org/10.1101/128728

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

The histone demethylase Kdm5 and the ARGONAUTE proteins Piwi and Aubergine regulate female abdominal pigmentation in Drosophila melanogaster

Insect pigmentation is an ecologically critical trait influencing many physiological processes. In Drosophila melanogaster, abdominal pigmentation is sexually dimorphic: males have fully pigmented posterior segments, while females exhibit a posterior melanin stripe. Pigmentation relies on the expression of pigmentation genes that encode enzymes involved in pigment synthesis. These genes are tightly regulated during pupal and young adult stages. To expand the gene regulatory network of pigmentation genes, we conducted an RNAi screen using the yellow-Gal4 driver, expressed during the pupal stage in abdominal epidermis. One of the candidates from this screen, Kdm5, encodes a histone demethylase erasing the H3K4me3 histone mark catalyzed by the histone methyl-transferase Trithorax (Trx). We show that Kdm5 down-regulation reduces abdominal pigmentation, mimicking trx down-regulation. Kdm5 activates melanin production through regulation of the pigmentation gene tan. Transcriptomic analyses reveal that Kdm5 and Trx share many targets in pupal abdominal epidermis, including piRNA pathway components such as piwi and aubergine. These piRNA components, originally associated with transposon silencing in the germline, also function in some somatic tissues such as the nervous system, the fat body or the gut. We demonstrate that Piwi and Aubergine participate in female abdominal pigmentation establishment, without evident piRNA production. We also show that Kdm5 and Piwi act not only in pupal abdominal epidermis but also in pupal fat body. This study therefore expands the regulatory network of pigmentation genes. It identifies a new somatic function for Kdm5 and Piwi and reveals a role for pupal fat body in female abdominal pigmentation regulation.

genetics

Genetic diversity within and between polyploid sugarcane (Saccharum spp.) families obtained via caryopsis using microsatellite markers and multicategory model

Genetic diversity analyses are essential for sugarcane (Saccharum spp.) breeding programs. Crossbreeding, based on genetic distances between parental plants, is a tool used to increase genetic variability and enhance plant selection; however, quantifying variation in highly polyploid species remains a challenge. The present study aimed to evaluate the diversity within and between 12 families of sugarcane derived from caryopses, analyzing 120 individual seedlings arranged in an augmented block design. Genotyping was performed using primers for 16 microsatellite loci, five simple sequence repeat (SSR) loci, and 11 expressed sequence tag-SSR (EST-SSR) loci. To accurately account for polyploidy, similarity calculations were performed using Bruvos distances among individuals and RST distances among the families. Analysis of molecular variance (AMOVA) indicated that most of the genetic variability was within families (72%), with only 28% found between them. This high level of intra-family variation demonstrates that a significant reservoir of genetic diversity remains available within the crosses. The highest genetic similarity was observed between the families RB986952 x RB986960 and RB036122 x RB03611, whereas the lowest genetic similarity was observed between the families RB97319 x RB966928 and RB106802 x RB855036. Although the evaluated families shared high genetic similarity, the pronounced genetic variation within them demonstrates a robust recombination potential, indicating that the genetic basis of sugarcane can be better explored using the high variability that already exists in the selection of desirable morpho-agronomic characteristics within the families. Furthermore, this study highlights the importance of using appropriate distances for diversity studies with codominant markers, such as microsatellites, in polyploid species.

genetics

Optimizing DNA extraction from environmentally degraded bone samples for molecular identification of cetacean species

Molecular identification of cetacean bone remains can be limited by DNA degradation and the presence of PCR inhibitors. Here, we present an optimized DNA extraction protocol based on a total demineralization method for environmentally exposed cetacean bones. The protocol uses 100 mg of bone powder, 24 h digestion with EDTA, N-lauroylsarcosine, and proteinase K, followed by a modified silica-column purification. Nine environmentally degraded bone samples representing eight individuals were processed. DNA concentrations ranged from 7.3 to 57.1 ng/uL (mean SD = 25.91- 13.91 ng/uL). The mitochondrial cytochrome b gene was successfully amplified from all samples using conventional PCR, and five samples (55.6%) yielded sequences suitable for downstream analysis. BLASTn identified Balaenoptera physalus as the closest database match for all recovered sequences, and phylogenetic analysis further supported their association with B. physalus reference sequences. These results demonstrate that the proposed protocol provides a practical approach for recovering amplifiable and molecularly informative mitochondrial DNA from environmentally degraded cetacean bone material, facilitating molecular identification from challenging skeletal remains.

genetics