Search bioRxivSearch

Biology subjects

Powell, J.

Publications and source records attributed to Powell, J..

12 recordsLinked to original sources

Unraveling the polygenic architecture of complex traits using blood eQTL meta-analysis

SummaryWhile many disease-associated variants have been identified through genome-wide association studies, their downstream molecular consequences remain unclear.\n\nTo identify these effects, we performed cis- and trans-expression quantitative trait locus (eQTL) analysis in blood from 31,684 individuals through the eQTLGen Consortium.\n\nWe observed that cis-eQTLs can be detected for 88% of the studied genes, but that they have a different genetic architecture compared to disease-associated variants, limiting our ability to use cis-eQTLs to pinpoint causal genes within susceptibility loci.\n\nIn contrast, trans-eQTLs (detected for 37% of 10,317 studied trait-associated variants) were more informative. Multiple unlinked variants, associated to the same complex trait, often converged on trans-genes that are known to play central roles in disease etiology.\n\nWe observed the same when ascertaining the effect of polygenic scores calculated for 1,263 genome-wide association study (GWAS) traits. Expression levels of 13% of the studied genes correlated with polygenic scores, and many resulting genes are known to drive these traits.

genomics

Generation of human neural retina transcriptome atlas by single cell RNA sequencing

The retina is a highly specialized neural tissue that senses light and initiates image processing. Although the functional organisation of specific cells within the retina has been well-studied, the molecular profile of many cell types remains unclear in humans. To comprehensively profile cell types in the human retina, we performed single cell RNA-sequencing on 20,009 cells obtained post-mortem from three donors and compiled a reference transcriptome atlas. Using unsupervised clustering analysis, we identified 18 transcriptionally distinct clusters representing all known retinal cells: rod photoreceptors, cone photoreceptors, Muller glia cells, bipolar cells, amacrine cells, retinal ganglion cells, horizontal cells, retinal astrocytes and microglia. Notably, our data captured molecular profiles for healthy and early degenerating rod photoreceptors, and revealed a novel role of MALAT1 in putative rod degeneration. We also demonstrated the use of this retina transcriptome atlas to benchmark pluripotent stem cell-derived cone photoreceptors and an adult Muller glia cell line. This work provides an important reference with unprecedented insights into the transcriptional landscape of human retinal cells, which is fundamental to our understanding of retinal biology and disease.

systems biology

Gene-Based Analysis in HRC Imputed Genome Wide Association Data Identifies Three Novel Genes For Alzheimer’s Disease

A novel POLARIS gene-based analysis approach was employed to compute gene-based polygenic risk score (PRS) for all individuals in the latest HRC imputed GERAD (N cases=3,332 and N controls=9,832) data using the International Genomics of Alzheimers Project summary statistics (N cases=13,676 and N controls=27,322, excluding GERAD subjects) to identify the SNPs and weight their risk alleles for the PRS score. SNPs were assigned to known, protein coding genes using GENCODE (v19). SNPs are assigned using both 1) no window around the gene and 2) a window of 35kb upstream and 10kb downstream to include transcriptional regulatory elements. The overall association of a gene is determined using a logistic regression model, adjusting for population covariates.\n\nThree novel gene-wide significant genes were determined from the POLARIS gene-based analysis using a gene window; PPARGC1A, RORA and ZNF423. The ZNF423 gene resides in an Alzheimers disease (AD)-specific protein network which also includes other AD-related genes. The PPARGC1A gene has been linked to energy metabolism and the generation of amyloid beta plaques and the RORA has strong links with genes which are differentially expressed in the hippocampus. We also demonstrate no enrichment for genes in either loss of function intolerant or conserved noncoding sequence regions.

genetics

A comprehensive assessment of benign genetic variability for neurodegenerative disorders

1AbstractOver the last few years, as more and more sequencing studies have been performed, it has become apparent that the identification of pathogenic mutations is, more often than not, a complex issue. Here, with a focus on neurodegenerative diseases, we have performed a survey of coding genetic variability that is unlikely to be pathogenic.\n\nWe have performed whole-exome sequencing in 478 samples derived from several brain banks in the United Kingdom and the United States of America. Samples were included when subjects were, at death, over 60 years of age, had no signs of neurological disease and were subjected to a neuropathological examination, which revealed no evidence of neurodegeneration. This information will be valuable to studies of genetic variability as a causal factor for neurodegenerative syndromes. We envisage it will be particularly relevant for diagnostic laboratories as a filter step to the results being produced by either genome-wide or gene-panel sequencing. We have made this data publicly available at www.alzforum.org/exomes/hex.

genomics

Detection of HPV E7 transcription at single-cell resolution in epidermis

Persistent human papillomavirus (HPV) infection is responsible for at least 5% of human malignancies. Most HPV-associated cancers are initiated by the HPV16 genotype, as confirmed by detection of integrated HPV DNA in cells of oral and anogenital epithelial cancers. However, single-cell RNA-sequencing (scRNA-seq) may enable prediction of HPV involvement in carcinogenesis at other sites. We conducted scRNA-seq on keratinocytes from a mouse transgenic for the E7 gene of HPV16, and showed sensitive and specific detection of HPV16-E7 mRNA, predominantly in basal keratinocytes. We showed that increased E7 mRNA copy number per cell was associated with increased expression of E7 induced genes. This technique enhances detection of viral transcripts in solid tissue and may clarify possible linkage of HPV infection to development of squamous cell carcinoma.

cell biology

Cardiac directed differentiation using small molecule Wnt modulation at single-cell resolution

Differentiation into diverse cell lineages requires the orchestration of gene regulatory networks guiding diverse cell fate choices. Utilizing human pluripotent stem cells, we measured expression dynamics of 17,718 genes from 43,168 cells across five time points over a thirty day time-course of in vitro cardiac-directed differentiation. Unsupervised clustering and lineage prediction algorithms were used to map fate choices and transcriptional networks underlying cardiac differentiation. We leveraged this resource to identify strategies for controlling in vitro differentiation as it occurs in vivo. HOPX, a non-DNA binding homeodomain protein essential for heart development in vivo was identified as dys-regulated in in vitro derived cardiomyocytes. Utilizing genetic gain and loss of function approaches, we dissect the transcriptional complexity of the HOPX locus and identify the requirement of hypertrophic signaling for HOPX transcription in hPSC-derived cardiomyocytes. This work provides a single cell dissection of the transcriptional landscape of cardiac differentiation for broad applications of stem cells in cardiovascular biology.

developmental biology

Determining cell fate specification and genetic contribution to cardiac disease risk in hiPSC-derived cardiomyocytes at single cell resolution

The majority of genetic loci underlying common disease risk act through changing genome regulation, and are routinely linked to expression quantitative trait loci, where gene expression is measured using bulk populations of mature cells. A crucial step that is missing is evidence of variation in the expression of these genes as cells progress from a pluripotent to mature state. This is especially important for cardiovascular disease, as the majority of cardiac cells have limited properties for renewal postneonatal. To investigate the dynamic changes in gene expression across the cardiac lineage, we generated RNA-sequencing data captured from 43,168 single cells progressing through in vitro cardiac-directed differentiation from pluripotency. We developed a novel and generalized unsupervised cell clustering approach and a machine learning method for prediction of cell transition. Using these methods, we were able to reconstruct the cell fate choices as cells transition from a pluripotent state to mature cardiomyocytes, uncovering intermediate cell populations that do not progress to maturity, and distinct cell trajectories that terminate in cardiomyocytes that differ in their contractile forces. Second, we identify new gene markers that denote lineage specification and demonstrate a substantial increase in their utility for cell identification over current pluripotent and cardiogenic markers. By integrating results from analysis of the single cell lineage RNA-sequence data with population-based GWAS of cardiovascular disease and cardiac tissue eQTLs, we show that the pathogenicity of disease-associated genes is highly dynamic as cells transition across their developmental lineage, and exhibit variation between cell fate trajectories. Through the integration of single cell RNA-sequence data with population-scale genetic data we have identified genes significantly altered at cell specification events providing insights into a context-dependent role in cardiovascular disease risk. This study provides a valuable data resource focused on in vitro cardiomyocyte differentiation to understand cardiac disease coupled with new analytical methods with broad applications to single-cell data.

genomics

ascend: R package for analysis of single cell RNA-seq data

Summaryascend is an R package comprised of fast, streamlined analysis functions optimized to address the statistical challenges of single cell RNA-seq. The package incorporates novel and established methods to provide a flexible framework to perform filtering, quality control, normalization, dimension reduction, clustering, differential expression and a wide-range of plotting. ascend is designed to work with scRNA-seq data generated by any high-throughput platform, and includes functions to convert data objects between software packages.\n\nAvailabilityThe R package and associated vignettes are freely available at https://github.com/IMB-Computational-Genomics-Lab/ascend.\n\nContactjoseph.powell@uq.edu.au\n\nSupplementary informationAn example dataset is available at ArrayExpress, accession number E-MTAB-6108

bioinformatics

Amplification-free, CRISPR-Cas9 Targeted Enrichment and SMRT Sequencing of Repeat-Expansion Disease Causative Genomic Regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats.\n\nWe have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtons disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.

genomics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Allelic Differentiation Of Complex Trait Loci Across Human Populations

How genetic variation contributes to phenotypic variation is a central question in genetics. Association signals for a complex trait are found throughout the majority of the genome suggesting much of the genome is under some degree of genetic constraint. Here, we develop a intraspecific population genetics approach to define a measure of population structure for each single nucleotide polymorphism (SNP). Using this approach, we test for evidence of stabilizing selection at complex traits and pleiotropic loci arising from the evolutionary history of 47 complex traits and common diseases. Our approach allowed us to identify traits and regions under stabilizing selection towards both global and subpopulation optima. Strongest depletion of allelic diversity was found at disease loci, indicating stabilizing selection has acted on these phenotypes in all subpopulations. Pleiotropic loci predominantly displayed evidence of stabilizing selection, often contributed to multiple disease risks, and sometimes also affected non-disease traits such as height. Risk alleles at pleiotropic disease loci displayed a more consistent direction of effect than expected by chance suggesting that stabilizing selection acting on pleiotropic loci is amplified through multiple disease phenotypes.

evolutionary biology

Single-Cell Transcriptome Sequencing Of 18,787 Human Induced Pluripotent Stem Cells Identifies Differentially Primed Subpopulations

Heterogeneity of cell states represented in pluripotent cultures have not been described at the transcriptional level. Since gene expression is highly heterogeneous between cells, single-cell RNA sequencing can be used to identify how individual pluripotent cells function. Here, we present results from the analysis of single-cell RNA sequencing data from 18,787 individual WTC CRISPRi human induced pluripotent stem cells. We developed an unsupervised clustering method, and through this identified four subpopulations distinguishable on the basis of their pluripotent state including: a core pluripotent population (48.3%), proliferative (47.8%), early-primed for differentiation (2.8%) and late-primed for differentiation (1.1%). For each subpopulation we were able to identify the genes and pathways that define differences in pluripotent cell states. Our method identified four transcriptionally distinct predictor gene sets comprised of 165 unique genes that denote the specific pluripotency states; and using these sets, we developed a multigenic machine learning prediction method to accurately classify single cells into each of the subpopulations. Compared against a set of established pluripotency markers, our method increases prediction accuracy by 10%, specificity by 20%, and explains a substantially larger proportion of deviance (up to 3-fold) from the prediction model. Finally, we developed an innovative method to predict cells transitioning between subpopulations, and support our conclusions with results from two orthogonal pseudotime trajectory methods.

genomics