Search bioRxivSearch

Biology subjects

Powell, J. E.

Publications and source records attributed to Powell, J. E..

5 recordsLinked to original sources

scPred: Single cell prediction using singular value decomposition and machine learning classification

Single-cell RNA sequencing has enabled the characterization of highly specific cell types in many human tissues, as well as both primary and stem cell-derived cell lines. An important facet of these studies is the ability to identify the transcriptional signatures that define a cell type or state. In theory, this information can be used to classify an unknown cell based on its transcriptional profile; and clearly, the ability to accurately predict a cell type and any pathologic-related state will play a critical role in the early diagnosis of disease and decisions around the personalized treatment for patients. Here we present a new generalizable method (scPred) for prediction of cell type(s), using a combination of unbiased feature selection from a reduced-dimension space, and machine-learning classification. scPred solves several problems associated with the identification of individual gene feature selection, and is able to capture subtle effects of many genes, increasing the overall variance explained by the model, and correspondingly improving the prediction accuracy. We validate the performance of scPred by performing experiments to classify tumor versus non-tumor epithelial cells in gastric cancer, then using independent molecular techniques (cyclic immunohistochemistry) to confirm our prediction, achieving an accuracy of classifying the disease state of individual cells of 99%. Moreover, we apply scPred to scRNA-seq data from pancreatic tissue, colorectal tumor biopsies, and circulating dendritic cells, and show that scPred is able to classify cell subtypes with an accuracy of 96.1-99.2%. Collectively, our results demonstrate the utility of scPred as a single cell prediction method that can be used for a wide variety of applications. The generalized method is implemented in software available here: https://github.com/IMB-Computational-Genomics-Lab/scPred/

genomics

Longitudinal expression profiling of CD4+ and CD8+ cells in patients with active to quiescent Giant Cell Arteritis

BackgroundGiant cell arteritis (GCA) is the most common form of vasculitis affecting elderly people. It is one of the few true ophthalmic emergencies. GCA is a heterogenous disease, symptoms and signs are variable thereby making it challenging to diagnose and often delaying diagnosis. A temporal artery biopsy is the gold standard to test for GCA, and there are currently no specific biochemical markers to categorize or aid diagnosis of the disease. We aimed to identify a less invasive method to confirm the diagnosis of GCA, as well as to ascertain clinically relevant predictive biomarkers by studying the transcriptome of purified peripheral CD4+ and CD8+ T lymphocytes in patients with GCA.\n\nMethods and FindingsWe recruited 16 patients with histological evidence of GCA at the Royal Victorian Eye and Ear Hospital (RVEEH), Melbourne, Australia, and aimed to collect blood samples at six time points: acute phase, 2-3 weeks, 6-8 weeks, 3 months, 6 months and 12 months after clinical diagnosis. CD4+ and CD8+ T-cells were positively selected at each time point through magnetic-assisted cell sorting (MACS). RNA was extracted from all 195 collected samples for subsequent RNA sequencing. The expression profiles of patients were compared to those of 16 age-matched controls. Over the 12-month study period, polynomial modelling analyses identified 179 and 4 statistically significant transcripts with altered expression profiles (FDR < 0.05) between cases and controls in CD4+ and CD8+ populations, respectively. In CD8+ cells, we identified two transcripts that remained differentially expressed after 12 months, namely SGTB, associated with neuronal apoptosis, and FCGR3A, which has been found in association with Takayasu arteritis (TA), another large vessel vasculitis. We detected genes that correlate with both symptoms and biochemical markers used in the acute setting for predicting long-term prognosis. 15 genes were shared across 3 phenotypes in CD4 and 16 across CD8 cells. In CD8, IL32 was common to 5 phenotypes: a history of Polymyalgia Rheumatica, both visual disturbance and raised neutrophils at the time of presentation, bilateral blindness and death within 12 months. Altered IL32 gene expression could provide risk evaluation of GCA diagnosis at the time of presentation and give an indication of prognosis, which may influence management.\n\nConclusionsThis is the first longitudinal gene expression study undertaken to identify robust transcriptomic biomarkers of GCA. Our results show cell type-specific transcript expression profiles, novel gene-phenotype associations, and uncover important biological pathways for this disease. These data significantly enhance the current knowledge of relevant biomarkers, their association with clinical prognostic markers, as well as potential candidates for detecting disease activity in whole blood samples. In the acute phase, the gene-phenotype relationships we have identified could provide insight to potential disease severity and as such guide us in initiating appropriate patient management.

genomics

Single Cell RNA Sequencing of stem cell-derived retinal ganglion cells.

We used human embryonic stem cell-derived retinal ganglion cells (RGCs) to characterize the transcriptome of 1,174 cells at the single cell level. The human embryonic stem cell line BRN3B-mCherry A81-H7 was differentiated to RGCs using a guided differentiation approach. Cells were harvested at day 36 and subsequently prepared for single cell RNA sequencing. Our data indicates the presence of three distinct subpopulations of cells, with various degrees of maturity. One cluster of 288 cells upregulated genes involved in axon guidance together with semaphorin interactions, cell-extracellular matrix interactions and ECM proteoglycans, suggestive of a more mature phenotype.

genomics

Identification of 55,000 Replicated DNA Methylation QTL

DNA methylation plays an important role in the regulation of transcription. Genetic control of DNA methylation is a potential candidate for explaining the many identified SNP associations with disease that are not found in coding regions. We replicated 52,916 cis and 2,025 trans DNA methylation quantitative trait loci (mQTL) using methylation measured on Illumina HumanMethylation450 arrays in the Brisbane Systems Genetics Study (n=614 from 177 families) and the Lothian Birth Cohorts of 1921 and 1936 (combined n = 1366). The trans mQTL SNPs were found to be over-represented in 1Mbp subtelomeric regions, and on chromosomes 16 and 19. There was a significant increase in trans mQTL DNA methylation sites in upstream and 5 UTR regions. No association was observed between either the SNPs or DNA methylation sites of trans mQTL and telomere length. The genetic heritability of a number of complex traits and diseases was partitioned into components due to mQTL and the remainder of the genome. Significant enrichment was observed for height (p = 2.1x10-10), ulcerative colitis (p = 2x10-5), Crohns disease (p = 6x10-8) and coronary artery disease (p = 5.5x10-6) when compared to a random sample of SNPs with matched minor allele frequency, although this enrichment is explained by the genomic location of the mQTL SNPs.

genomics

Constraints on eQTL fine mapping in the presence of multi-site local regulation of gene expression

Expression QTL (eQTL) detection has emerged as an important tool for unravelling of the relationship between genetic risk factors and disease or clinical phenotypes. Most studies use single marker linear regression to discover primary signals, followed by sequential conditional modeling to detect secondary genetic variants affecting gene expression. However, this approach assumes that functional variants are sparsely distributed and that close linkage between them has little impact on estimation of their precise location and magnitude of effects. In this study, we address the prevalence of secondary signals and bias in estimation of their effects by performing multi-site linear regression on two large human cohort peripheral blood gene expression datasets (each greater than 2,500 samples) with accompanying whole genome genotypes, namely the CAGE compendium of Illumina microarray studies, and the Framingham Heart Study Affymetrix data. Stepwise conditional modeling demonstrates that multiple eQTL signals are present for ~40% of over 3500 eGenes in both datasets, and the number of loci with additional signals reduces by approximately two-thirds with each conditioning step. However, the concordance of specific signals between the two studies is only ~30%, indicating that expression profiling platform is a large source of variance in effect estimation. Furthermore, a series of simulation studies imply that in the presence of multi-site regulation, up to 10% of the secondary signals could be artefacts of incomplete tagging, and at least 5% but up to one quarter of credible intervals may not even include the causal site, which is thus mis-localized. Joint multi-site effect estimation recalibrates effect size estimates by just a small amount on average. Presumably similar conclusions apply to most types of quantitative trait. Given the strong empirical evidence that gene expression is commonly regulated by more than one variant, we conclude that the fine-mapping of causal variants needs to be adjusted for multi-site influences, as conditional estimates can be highly biased by interference among linked sites.

genetics