Search bioRxiv⌕ Search

Biology subjects

Gordon, M. G.

Publications and source records attributed to Gordon, M. G..

4 recordsLinked to original sources

Tensor decomposition reveals coordinated multicellular patterns of transcriptional variation that distinguish and stratify disease individuals

Tissue- and organism-level biological processes often involve coordinated action of multiple distinct cell types. Current computational methods for the analysis of single-cell RNA-sequencing (scRNA-seq) data, however, are not designed to capture co-variation of cell states across samples, in part due to the low number of biological samples in most scRNA-seq datasets. Recent advances in sample multiplexing have enabled population-scale scRNA-seq measurements of tens to hundreds of samples. To take advantage of such datasets, here we introduce a computational approach called single-cell Interpretable Tensor Decomposition (scITD). This method extracts "multicellular gene expression patterns" that capture how sample-specific expression states of a cell type are correlated with the expression states of other cell types. Such multicellular patterns can reveal molecular mechanisms underlying coordinated changes of different cell types within the tissue, and can be used to stratify individuals in a clinically-relevant and reproducible manner. We first validated the performance of scITD using in vitro experimental data and simulations. We then applied scITD to scRNA-seq data on peripheral blood mononuclear cells (PBMCs) from 115 patients with systemic lupus erythematosus and 56 healthy controls. We recapitulated a well-established pan-cell-type signature of interferon-signaling that was associated with the presence of anti-dsDNA autoantibodies and a disease activity index. We further identified a novel multicellular pattern linked to nephritis, which was characterized by an expansion of activated memory B cells along with helper T cell activation. Our approach also sheds light on ligand-receptor interactions potentially mediating these multicellular patterns. As validation, we demonstrated that these expression patterns also stratified donors from a pediatric SLE dataset by the same phenotypic attributes. Lastly, we found the interferon multicellular pattern and others to be conserved in a COVID-19 dataset, pointing to the presence of both general and disease-specific patterns of inter-individual immune variation. Overall, scITD is a flexible method for exploring co-variation of cell states in multi-sample single-cell datasets, which can yield new insights into complex non-cell-autonomous dependencies that define and stratify disease.

bioinformatics↗

Multi-context genetic modeling of transcriptional regulation resolves novel disease loci

A majority of the variants identified in genome-wide association studies fall in non-coding regions of the genome, indicating their mechanism of impact is mediated via gene expression. Leveraging this hypothesis, transcriptome-wide association studies (TWAS) have assisted in both the interpretation and discovery of additional genes associated with complex traits. However, existing methods for conducting TWAS do not take full advantage of the intra-individual correlation inherently present in multi-context expression studies and do not properly adjust for multiple testing across contexts. We developed CONTENT-- a computationally efficient method with proper cross-context false discovery correction that leverages correlation structure across contexts to improve power and generate context-specific and context-shared components of expression. We applied CONTENT to bulk multi-tissue and single-cell RNA-seq data sets and show that CONTENT leads to a 42% (bulk) and 110% (single cell) increase in the number of genetically predicted genes relative to previous approaches. Interestingly, we find the context-specific component of expression comprises 30% of heritability in tissue-level bulk data and 75% in single-cell data, consistent with cell type heterogeneity in bulk tissue. In the context of TWAS, CONTENT increased the number of gene-phenotype associations discovered by over 47% relative to previous methods across 22 complex traits.

genomics↗

Fast and powerful statistical method for context-specific QTL mapping in multi-context genomic studies

Context-specific eQTLs mediate genetic risk for complex diseases. However, limitations in current methods for identifying these eQTLs have hindered their comprehensive characterization and downstream interpretation of disease-associated variants. Here, we introduce FastGxC, a method to efficiently and powerfully map context-specific eQTLs by leveraging the correlation structure in genomic studies with repeated sampling, e.g., single-cell RNA-seq studies. Using simulations, we demonstrate that FastGxC is up to nine times more powerful and 106 times faster than existing approaches, reducing computation time from years to minutes. We applied FastGxC to bulk multi-tissue (N=698) and single-cell PBMC (N=1,218) RNA-seq datasets, generating comprehensive tissue- and cell-type-specific eQTL maps. These eQTLs exhibited up to four-fold enrichment in open chromatin regions from matched contexts and were twice as enriched as standard context-specific eQTLs, highlighting their biological relevance. Furthermore, we examined the relationship between context-specific eQTLs and complex human traits and diseases. FastGxC improved precision in identifying relevant contexts for each trait by three-fold and expanded candidate causal genes by 25% in cell types and 6% in tissues compared to standard eQTLs. In summary, FastGxC provides a powerful framework for mapping context-specific eQTLs, advancing our understanding of gene regulatory mechanisms underlying complex human traits and diseases.

genomics↗

Robust Sequence Determinants of α-Synuclein Toxicity in Yeast Implicate Membrane Binding

Protein conformations are shaped by cellular environments, but how environmental changes alter the conformational landscapes of specific proteins in vivo remains largely uncharacterized, in part due to the challenge of probing protein structures in living cells. Here, we use deep mutational scanning to investigate how a toxic conformation of -synuclein, a dynamic protein linked to Parkinsons disease, responds to perturbations of cellular proteostasis. In the context of a course for graduate students in the UCSF Integrative Program in Quantitative Biology, we screened a comprehensive library of -synuclein missense mutants in yeast cells treated with a variety of small molecules that perturb cellular processes linked to -synuclein biology and pathobiology. We found that the conformation of -synuclein previously shown to drive yeast toxicity--an extended, membrane-bound helix--is largely unaffected by these chemical perturbations, underscoring the importance of this conformational state as a driver of cellular toxicity. On the other hand, the chemical perturbations have a significant effect on the ability of mutations to suppress -synuclein toxicity. Moreover, we find that sequence determinants of -synuclein toxicity are well described by a simple structural model of the membrane-bound helix. This model predicts that -synuclein penetrates the membrane to constant depth across its length but that membrane affinity decreases toward the C terminus, which is consistent with orthogonal biophysical measurements. Finally, we discuss how parallelized chemical genetics experiments can provide a robust framework for inquiry-based graduate coursework.

biochemistry↗