Search bioRxiv⌕ Search

Biology subjects

Sverchkov, Y.

Publications and source records attributed to Sverchkov, Y..

3 recordsLinked to original sources

A graph-based learning approach to predict the effects of gene perturbations on molecular phenotypes

MotivationLarge-scale gene knockdown/knockout screens have been used to gain insight into a wide array of phenotypes and biological processes. However, conducting such experiments is expensive and labor-intensive. In this work, we present a general graph-based machine-learning approach that can predict the effects of gene perturbations on molecular phenotypes of interest given some measured phenotypic effects of other gene perturbations. The motivation for learning models that can predict the effects of gene perturbations is fourfold. Such models can (1) predict effects for unmeasured genes in cases in which cost or technical barriers preclude perturbing every gene, (2) prioritize unmeasured genes or sets of genes for subsequent perturbation experiments, (3) hypothesize mechanisms that underlie the relationships between the perturbed genes and their effects, and (4) generalize to other unmeasured phenotypes of interest. ResultsWe evaluate our approach by applying it, in conjunction with four different learning methods, to learn models for four varied phenotypes. Our empirical evaluation demonstrates that the learned models (1) show relatively high levels of predictive accuracy across the four phenotypes, (2) have better predictive accuracy than several standard baselines, (3) can often learn accurate models with small training sets, (4) benefit from having multiple sources of evidence in the input representation, (5) can, in many cases, transfer their predictive value to other phenotypes. Data availabilityThe assembled data sets and source code for this work are available at: https://github.com/Craven-Biostat-Lab/graph-molecular-phenotype-prediction Author summaryOne general approach for gaining insight into the genes involved in a specific biological process is to conduct an experiment in which individual genes are perturbed and the effect on the process is measured for each perturbation. Large-scale experiments of this type have provided important biological insights, but they are often expensive and labor-intensive to perform. As a result, it is not always feasible to measure the effects of perturbing every gene. In this article, we present a machine-learning approach to predicting the effects of gene perturbations using available experimental data and biological network information. Our method can estimate the effects of genes that have not yet been experimentally measured, helping researchers identify promising genes to study next. In addition, the models can suggest hypotheses about the molecular interactions that link genes to the biological process of interest. Approaches like this may help guide experimental studies and accelerate the discovery of gene-phenotype relationships.

systems biology↗

Gene- and domain-aware calibration increases the clinical utility of variant effect predictors

The utility of clinical genetic testing is limited because around 90% of missense variants in ClinVar remain of uncertain clinical significance. Variant effect predictors (VEPs) can score any missense variant, potentially empowering variant classification. Realizing this potential requires calibration to translate predictor scores into evidence. However, genome-wide calibration ignores predictor heterogeneity across genes, causing evidence misassignment. We developed an automated, data-adaptive framework that optimizes two complementary approaches: gene-specific calibration for genes with enough variants for calibration, and domain-aggregate calibration for other disease-associated genes, which groups variants from protein domains with similar predictor score distributions for calibration. Applied to three predictors across 2,769 genes, this framework assigned evidence to 10.6% more variants on average while generally improving evidence accuracy compared to genome-wide calibration. These calibrations and the resulting calibrated computational evidence are available through the PredictMD portal. Our framework substantially increases the clinical utility of VEPs for variant classification.

genetics↗

A scalable approach to resolving variants of uncertain significance

Over 90% of missense variants across [~]4,000 disease-associated genes are variants of uncertain significance (VUS). Experimental variant effect measurements provide critical evidence about pathogenicity and inform disease biology, but most variants lack data and clinical translation has been limited. The Impact of Genomic Variation on Function Consortium generated experimental data for 62,215 variants across ten genes using multiplexed assays and 1,407 variants across 163 genes using arrayed assays, curated 193,139 additional community-generated variant effect measurements across 30 additional genes, and developed automated calibration methods for translating experimental data and variant effect predictions into clinical evidence. To reduce current VUS, we developed a scalable workflow using only experimental and predictive evidence, enabling reclassification of 75% of the 16,115 VUS in these genes as pathogenic or benign with <1% error. To minimize future VUS, we analyzed >90,000 unobserved variants; 62% had enough evidence to be "preclassified" as pathogenic or benign. We validated our data, evidence and classifications using All of Us and created interactive resources to enable clinical use of the calibrated data. Thus, for 40 genes, representing 1% of the clinical genome, we resolve most existing VUS and future variants, illustrating how systematic use of scalable evidence can empower genomic medicine.

genomics↗