Search bioRxiv⌕ Search

Biology subjects

Gould, S. I.

Publications and source records attributed to Gould, S. I..

3 recordsLinked to original sources

Multiplexed in vivo base editing identifies functional gene-variant-context interactions

Human genome sequencing efforts in healthy and diseased individuals continue to identify a broad spectrum of genetic variants associated with predisposition, progression, and therapeutic outcomes for diseases like cancer1-6. Insights derived from these studies have significant potential to guide clinical diagnoses and treatment decisions; however, the relative importance and functional impact of most genetic variants remain poorly understood. Precision genome editing technologies like base and prime editing can be used to systematically engineer and interrogate diverse types of endogenous genetic variants in their native context7-9. We and others have recently developed and applied scalable sensor-based screening approaches to engineer and measure the phenotypes produced by thousands of endogenous mutations in vitro10-12. However, the impact of most genetic variants in the physiological in vivo setting, including contextual differences depending on the tissue or microenvironment, remains unexplored. Here, we integrate new cross-species base editing sensor libraries with syngeneic cancer mouse models to develop a multiplexed in vivo platform for systematic functional analysis of endogenous genetic variants in primary and disseminated malignancies. We used this platform to screen 13,840 guide RNAs designed to engineer 7,783 human cancer-associated mutations mapping to 489 endogenous protein-coding genes, allowing us to construct a rich compendium of putative functional interactions between genes, mutations, and physiological contexts. Our findings suggest that the physiological in vivo environment and cellular organotropism are important contextual determinants of specific gene-variant phenotypes. We also show that many mutations and their in vivo effects fail to be detected with standard CRISPR-Cas9 nuclease approaches and often produce discordant phenotypes, potentially due to site-specific amino acid selection- or separation-of-function mechanisms. This versatile platform could be deployed to investigate how genetic variation impacts diverse in vivo phenotypes associated with cancer and other genetic diseases, as well as identify new potential therapeutic avenues to treat human disease.

cancer biology↗

Computational modeling of human genetic variants in mice

Mouse models represent a powerful platform to study genes and variants associated with human diseases. While genome editing technologies have increased the rate and precision of model development, predicting and installing specific types of mutations in mice that mimic the native human genetic context is complicated. Computational tools can identify and align orthologous wild-type genetic sequences from different species; however, predictive modeling and engineering of equivalent mouse variants that mirror the nucleotide and/or polypeptide change effects of human variants remains challenging. Here, we present H2M (human-to-mouse), a computational pipeline to analyze human genetic variation data to systematically model and predict the functional consequences of equivalent mouse variants. We show that H2M can integrate mouse-to-human and paralog-to-paralog variant mapping analyses with precision genome editing pipelines to devise strategies tailored to model specific variants in mice. We leveraged these analyses to establish a database containing > 3 million human-mouse equivalent mutation pairs, as well as in silico-designed base and prime editing libraries to engineer 4,944 recurrent variant pairs. Using H2M, we also found that predicted pathogenicity and immunogenicity scores were highly correlated between human-mouse variant pairs, suggesting that variants with similar sequence change effects may also exhibit broad interspecies functional conservation. Overall, H2M fills a gap in the field by establishing a robust and versatile computational framework to identify and model homologous variants across species while providing key experimental resources to augment functional genetics and precision medicine applications. The H2M database (including software package and documentation) can be accessed at https://human2mouse.com.

bioinformatics↗

PEGG: A computational pipeline for rapid design of prime editing guide RNAs and sensor libraries

Many human diseases have a strong association with diverse types of genetic alterations. These diseases include cancer, in which tumor genomes often harbor a complex spectrum of single-nucleotide alterations and chromosomal rearrangements that can perturb gene function in ways that remain poorly understood. Some cancer-associated genes exhibit a tremendous degree of mutational heterogeneity, which may impact disease initiation, progression, and therapy responses. For example, TP53, the most frequently mutated gene in cancer, shows extensive allelic variation that leads to the generation of altered proteins that can produce functionally distinct phenotypes. Whether distinct variants of TP53 and other genes encode proteins with loss-of-function, gain-of-function, or otherwise neomorphic phenotypes remains both controversial and technically challenging to assess, particularly at the endogenous level. Here, we present a high-throughput prime editing "sensor" strategy to quantitatively assess the functional impact of diverse types of endogenous genetic variants. We used this strategy to screen the largest collection of endogenous cancer-associated TP53 variants assembled to date, identifying both known and novel alleles that impact p53 function in mechanistically diverse ways. Intriguingly, we find that certain types of endogenous TP53 variants, particularly those in the p53 oligomerization domain, display opposite phenotypes in exogenous overexpression systems. These include disease-relevant variants found in humans with cancer predisposition syndromes that encode altered proteins with unique molecular properties. Our results emphasize the physiological importance of gene dosage in shaping native protein stoichiometry and protein-protein interactions, highlight the dangers of using exogenous overexpression systems to interpret pathogenic alleles, and establish a powerful computational and experimental framework for studying diverse types of genetic variants in their endogenous sequence context at scale.

bioinformatics↗