Search bioRxivSearch

Biology subjects

Cassa, C.

Publications and source records attributed to Cassa, C..

2 recordsLinked to original sources

Applicability of the mutation-selection balance model to population genetics of heterozygous protein-truncating variants in humans

The fate of alleles in the human population is believed to be highly affected by the stochastic force of genetic drift. Estimation of the strength of natural selection in humans generally necessitates a careful modeling of drift including complex effects of the population history and structure. Protein truncating variants (PTVs) are expected to evolve under strong purifying selection and to have a relatively high per-gene mutation rate. Thus, it is appealing to model the population genetics of PTVs under a simple deterministic mutation-selection balance, as has been proposed earlier [1]. Here, we investigated the limits of this approximation using both computer simulations and data-driven approaches. Our simulations rely on a model of demographic history estimated from 33,370 individual exomes of the Non-Finnish European subset of the ExAC dataset [2]. Additionally, we compared the African and European subset of the ExAC study and analyzed de novo PTVs. We show that the mutation-selection balance model is applicable to the majority of human genes, but not to genes under the weakest selection.

genetics

A literature review at genome scale: improving clinical variant assessment

Introduction: Over 150,000 variants have been reported to cause Mendelian disease in the medical literature. It is still difficult to leverage this knowledge base in clinical practice as many reports lack strong statistical evidence or may include false associations. Clinical laboratories assess whether these variants (along with newly observed variants that are adjacent to these published ones) underlie clinical disorders.\n\nMaterials and Methods: We measured whether citation data--including journal impact factor and the number of cited variants (NCV) in each gene with published disease associations--can be used to improve variant assessment.\n\nResults: Surprisingly, we find that impact factor is not predictive of pathogenicity, but the NCV score for each gene can provide statistical support of pathogenicity. When combining this gene-level citation metric with variant-level evolutionary conservation and structural features, classification accuracy reaches 89.5%. Further, variants identified in clinical exome sequencing cases have higher NCV scores than simulated rare variants from ExAC in matched genes and consequences (p<2.22x10-16).\n\nDiscussion: Aggregate citation data can complement existing variant-based predictive algorithms, and can boost their performance without accessing and reviewing large numbers of manuscripts. The NCV is a slow-growing metric of scientific knowledge about each genes association with disease.\n\nFunding: This research was supported by NIH NHGRI grant HG007229 (C.C.) and NIGMS grant GM078598 (I.A., D.J., and S.S.).

bioinformatics