Search bioRxivSearch

Biology subjects

Sunyaev, S.

Publications and source records attributed to Sunyaev, S..

3 recordsLinked to original sources

Genome wide association with quantitative resistance phenotypes in Mycobacterium tuberculosis reveals novel resistance genes and regulatory regions

Drug resistance is threatening attempts at tuberculosis epidemic control. Molecular diagnostics for drug resistance that rely on the detection of resistance-related mutations could expedite patient care and accelerate progress in TB eradication. We performed minimum inhibitory concentration testing for 12 anti-TB drugs together with Illumina whole genome sequencing on 1452 clinical Mycobacterium tuberculosis (MTB) isolates. We then used a linear mixed model to evaluate genome wide associations between mutations in MTB genes or noncoding regions and drug resistance, followed by validation of our findings in an independent dataset of 792 patient isolates. Novel associations at 13 genomic loci were confirmed in the validation set, with 2 involving noncoding regions. We found promoter mutations to have smaller average effects on resistance levels than gene body mutations in genes where both can contribute to resistance. Enabled by a quantitative measure of resistance, we estimated the heritability of the resistance phenotype to 11 anti-TB drugs and identify a lower than expected contribution from known resistance genes. We also report the proportion of variation in resistance levels explained by the novel loci identified here. This study highlights the complexity of the genomic mechanisms associated with the MTB resistance phenotype, including the relatively large number of potentially causative or compensatory loci, and emphasizes the contribution of the noncoding portion of the genome.

evolutionary biology

Low-frequency variant functional architectures reveal strength of negative selection across coding and non-coding annotations

Common variant heritability is known to be concentrated in variants within cell-type-specific non-coding functional annotations, with a limited role for common coding variants. However, little is known about the functional distribution of low-frequency variant heritability. Here, we partitioned the heritability of both low-frequency (0.5% [&le;] MAF < 5%) and common (MAF [&ge;] 5%) variants in 40 UK Biobank traits (average N = 363K) across a broad set of coding and non-coding functional annotations, employing an extension of stratified LD score regression to low-frequency variants that produces robust results in simulations. We determined that non-synonymous coding variants explain 17{+/-}1% of low-frequency variant heritability [Formula] versus only 2.1{+/-}0.2% of common variant heritability [Formula], and that regions conserved in primates explain nearly half of [Formula] (43{+/-}2%). Other annotations previously linked to negative selection, including non-synonymous variants with high PolyPhen-2 scores, non-synonymous variants in genes under strong selection, and low-LD variants, were also significantly more enriched for [Formula] as compared to [Formula]. Cell-type-specific non-coding annotations that were significantly enriched for [Formula] of corresponding traits tended to be similarly enriched for [Formula] for most traits, but more enriched for brain-related annotations and traits. For example, H3K4me3 marks in brain DPFC explain 57{+/-}12% of [Formula] vs. 12{+/-}2% of [Formula] for neuroticism, implicating the action of negative selection on low-frequency variants affecting gene regulation in the brain. Forward simulations confirmed that the ratio of low-frequency variant enrichment vs. common variant enrichment primarily depends on the mean selection coefficient of causal variants in the annotation, and can be used to predict the effect size variance of causal rare variants (MAF < 0.5%) in the annotation, informing their prioritization in whole-genome sequencing studies. Our results provide a deeper understanding of low-frequency variant functional architectures and guidelines for the design of association studies targeting functional classes of low-frequency and rare variants.

genetics

A literature review at genome scale: improving clinical variant assessment

Introduction: Over 150,000 variants have been reported to cause Mendelian disease in the medical literature. It is still difficult to leverage this knowledge base in clinical practice as many reports lack strong statistical evidence or may include false associations. Clinical laboratories assess whether these variants (along with newly observed variants that are adjacent to these published ones) underlie clinical disorders.\n\nMaterials and Methods: We measured whether citation data--including journal impact factor and the number of cited variants (NCV) in each gene with published disease associations--can be used to improve variant assessment.\n\nResults: Surprisingly, we find that impact factor is not predictive of pathogenicity, but the NCV score for each gene can provide statistical support of pathogenicity. When combining this gene-level citation metric with variant-level evolutionary conservation and structural features, classification accuracy reaches 89.5%. Further, variants identified in clinical exome sequencing cases have higher NCV scores than simulated rare variants from ExAC in matched genes and consequences (p<2.22x10-16).\n\nDiscussion: Aggregate citation data can complement existing variant-based predictive algorithms, and can boost their performance without accessing and reviewing large numbers of manuscripts. The NCV is a slow-growing metric of scientific knowledge about each genes association with disease.\n\nFunding: This research was supported by NIH NHGRI grant HG007229 (C.C.) and NIGMS grant GM078598 (I.A., D.J., and S.S.).

bioinformatics