Search bioRxivSearch

Biology subjects

Blangero, J.

Publications and source records attributed to Blangero, J..

8 recordsLinked to original sources

Efficient variant set mixed model association tests for continuous and binary traits in large-scale whole genome sequencing studies

With advances in Whole Genome Sequencing (WGS) technology, more advanced statistical methods for testing genetic association with rare variants are being developed. Methods in which variants are grouped for analysis are also known as variant-set, gene-based, and aggregate unit tests. The burden test and Sequence Kernel Association Test (SKAT) are two widely used variant-set tests, which were originally developed for samples of unrelated individuals and later have been extended to family data with known pedigree structures. However, computationally-efficient and powerful variant-set tests are needed to make analyses tractable in large-scale WGS studies with complex study samples. In this paper, we propose the variant-Set Mixed Model Association Tests (SMMAT) for continuous and binary traits using the generalized linear mixed model framework. These tests can be applied to large-scale WGS studies involving samples with population structure and relatedness, such as in the National Heart, Lung, and Blood Institutes Trans-Omics for Precision Medicine (TOPMed) program. SMMAT tests share the same null model for different variant sets, and a virtue of this null model, which includes covariates only, is that it needs to be only fit once for all tests in each genome-wide analysis. Simulation studies show that all the proposed SMMAT tests correctly control type I error rates for both continuous and binary traits in the presence of population structure and relatedness. We also illustrate our tests in a real data example of analysis of plasma fibrinogen levels in the TOPMed program (n = 23,763), using the Analysis Commons, a cloud-based computing platform.

genetics

Inferring identical by descent sharing of sample ancestors promotes high resolution relative detection

As genetic datasets increase in size, the fraction of samples with one or more close relatives grows rapidly, resulting in sets of mutually related individuals. We present DRUID--Deep Relatedness Utilizing Identity by Descent--a method that works by inferring the identical by descent (IBD) sharing profile of an ungenotyped ancestor of a set of close relatives. Using this IBD profile, DRUID infers relatedness between unobserved ancestors and more distant relatives, thereby combining information from multiple samples to remove one or more generations between the deep relationships to be identified. DRUID constructs sets of close relatives by detecting full siblings and also uses a novel approach to identify the aunts/uncles of two or more siblings, recovering 92.2% of real aunts/uncles with zero false positives. In real and simulated data, DRUID correctly infers up to 10.5% more relatives than PADRE when using data from two sets of distantly related siblings, and 10.7-31.3% more relatives given two sets of siblings and their aunts/uncles. DRUID frequently infers relationships either correctly or within one degree of the truth, with PADRE classifying 43.3-58.3% of tenth degree relatives in this way compared to 79.6-96.7% using DRUID.

genetics

Fast and Powerful Genome Wide Association Analysis of Dense Genetic Data with High Dimensional Imaging Phenotypes

Genome wide association (GWA) analysis of brain imaging phenotypes can advance our understanding of the genetic basis of normal and disorder-related variation in the brain. GWA approaches typically use linear mixed effect models to account for non-independence amongst subjects due to factors such as family relatedness and population structure. The use of these models with high-dimensional imaging phenotypes presents enormous challenges in terms of computational intensity and the need to account multiple testing in both the imaging and genetic domain. Here we present method that makes mixed models practical with high-dimensional traits by a combination of a transformation applied to the data and model, and the use of a non-iterative variance component estimator. With such speed enhancements permutation tests are feasible, which allows inference on powerful spatial tests like the cluster size statistic.

genetics

Whole Genome Sequencing in Psychiatric Disorders: the WGSPD Consortium

As technology advances, whole genome sequencing (WGS) is likely to supersede other genotyping technologies. The rate of this change depends on its relative cost and utility. Variants identified uniquely through WGS may reveal novel biological pathways underlying complex disorders and provide high-resolution insight into when, where, and in which cell type these pathways are affected. Alternatively, cheaper and less computationally intensive approaches may yield equivalent insights. Understanding the role of rare variants in the noncoding gene-regulating genome, through pilot WGS projects, will be critical to determine which of these two extremes best represents reality. With large cohorts, well-defined risk loci, and a compelling need to understand the underlying biology, psychiatric disorders have a role to play in this preliminary WGS assessment. The WGSPD consortium will integrate data for 18,000 individuals with psychiatric disorders, beginning with autism spectrum disorder, schizophrenia, bipolar disorder, and major depressive disorder, along with over 150,000 controls.

genomics

Accurate Phasing of Pedigree Genotypes Using Whole Genome Sequence Data

Phasing, the process of predicting haplotypes from genotype data, is an important undertaking in genetics and an ongoing area of research. Phasing methods, and associated software, designed specifically for pedigrees are urgently needed. Here we present a new method for phasing genotypes from whole genome sequencing data in pedigrees: PULSAR (Phasing Using Lineage Specific Alleles / Rare variants). The method is built upon the idea that alleles that are specific to a single founding chromosome within a pedigree, which we refer to as lineage-specific alleles, are highly informative for identifying haplotypes that are identical-by-decent between individuals within a pedigree. Through extensive simulation we assess the performance of PULSAR in a variety of pedigree sizes and structures, and we explore the effects of genotyping errors and presence of non-sequenced individuals on its performance. If the genotyping error rate is sufficiently low PULSAR can phase > 99.9% of heterozygous genotypes with a switch error rate below 1 x 10-4 in pedigrees where all individuals are sequenced. We demonstrate that the method is highly accurate and consistently outperforms the long-range phasing approach used for comparison in our benchmarking. The method also holds promise for fixing genotype errors or imputing missing genotypes. The software implementation of this method is freely available.

genetics

Do Candidate Genes Affect the Brain’s White Matter Microstructure? Large-Scale Evaluation of 6,165 Diffusion MRI Scans

AbstractSusceptibility genes for psychiatric and neurological disorders - including APOE, BDNF, CLU,CNTNAP2, COMT, DISC1, DTNBP1, ErbB4, HFE, NRG1, NTKR3, and ZNF804A - have been reported to affect white matter (WM) microstructure in the healthy human brain, as assessed through diffusion tensor imaging (DTI). However, effects of single nucleotide polymorphisms (SNPs) in these genes explain only a small fraction of the overall variance and are challenging to detect reliably in single cohort studies. To date, few studies have evaluated the reproducibility of these results. As part of the ENIGMA-DTI consortium, we pooled regional fractional anisotropy (FA) measures for 6,165 subjects (CEU ancestry N=4,458) from 11 cohorts worldwide to evaluate effects of 15 candidate SNPs by examining their associations with WM microstructure. Additive association tests were conducted for each SNP. We used several meta-analytic and mega-analytic designs, and we evaluated regions of interest at multiple granularity levels. The ENIGMA-DTI protocol was able to detect single-cohort findings as originally reported. Even so, in this very large sample, no significant associations remained after multiple-testing correction for the 15 SNPs investigated. Suggestive associations (1.3x10-4 < p < 0.05, uncorrected) were found for BDNF, COMT, and ZNF804A in specific tracts. Meta-and mega-analyses revealed similar findings. Regardless of the approach, the previously reported candidate SNPs did not show significant associations with WM microstructure in this largest genetic study of DTI to date; the negative findings are likely not due to insufficient power. Genome-wide studies, involving large-scale meta-analyses, may help to discover SNPs robustly influencing WM microstructure.

neuroscience

A performance assessment of relatedness inference methods using genome-wide data from thousands of relatives

Inferring relatedness from genomic data is an essential component of genetic association studies, population genetics, forensics, and genealogy. While numerous methods exist for inferring relatedness, thorough evaluation of these approaches in real data has been lacking. Here, we report an assessment of 12 state-of-the-art pairwise relatedness inference methods using a dataset with 2,485 individuals contained in several large pedigrees that span up to six generations. We find that all methods have high accuracy (~92% - 99%) when detecting first and second degree relationships, but their accuracy dwindles to less than 43% for seventh degree relationships. However, most IBD segment-based methods inferred seventh degree relatives correct to within one relatedness degree for more than 76% of relative pairs. Overall, the most accurate methods are ERSA and approaches that compute total IBD sharing using the output from GERMLINE and Refined IBD to infer relatedness. Combining information from the most accurate methods provides little accuracy improvement, indicating that novel approaches--such as new methods that leverage relatedness signals from multiple samples--are needed to achieve a sizeable jump in performance.

genetics

Genetic variation and gene expression across multiple tissues and developmental stages in a non-human primate

By analyzing multi-tissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalogue of expression quantitative trait loci (eQTLs) in a non-human primate model. This catalogue contains more genome-wide significant eQTLs, per sample, than comparable human resources, and reveals sex and age-related expression patterns. Findings include a master regulatory locus that likely plays a role in immune function, and a locus regulating hippocampal long non-coding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.

genetics