Search bioRxivSearch

Biology subjects

Daetwyler, H. D.

Publications and source records attributed to Daetwyler, H. D..

3 recordsLinked to original sources

Genome variants altering gene splicing in bovine are extensively shared between tissues

BackgroundMammalian phenotypes are shaped by numerous genome variants, many of which may regulate gene transcription or RNA splicing. To identify variants with regulatory functions in cattle, an important economic and model species, we used sequence variants to map a type of expression quantitative trait loci (expression QTLs) that are associated with variations in the RNA splicing, i.e., sQTLs. To further the understanding of regulatory variants, sQTLs were compare with other two types of expression QTLs, 1) variants associated with variations in gene expression, i.e., geQTLs and 2) variants associated with variations in exon expression, i.e., eeQTLs, in different tissues.\n\nResultsUsing whole genome and RNA sequence data from four tissues of over 200 cattle, sQTLs identified using exon inclusion ratios were verified by matching their effects on adjacent intron excision ratios. sQTLs contained the highest percentage of variants that are within the intronic region of genes and contained the lowest percentage of variants that are within intergenic regions, compared to eeQTLs and geQTLs. Many geQTLs and sQTLs are also detected as eeQTLs. Many expression QTLs, including sQTLs, were significant in all four tissues and had a similar effect in each tissue. To verify such expression QTL sharing between tissues, variants surrounding ({+/-}1Mb) the exon or gene were used to build local genomic relationship matrices (LGRM) and estimated genetic correlations between tissues. For many exons, the splicing and expression level was determined by the same cis additive genetic variance in different tissues. Thus, an effective but simple-to-implement meta-analysis combining information from three tissues is introduced to increase power to detect and validate sQTLs. sQTLs and eeQTLs together were more enriched for variants associated with cattle complex traits, compared to geQTLs. Several putative causal mutations were identified, including an sQTL at Chr6:87392580 within the 5th exon of kappa casein (CSN3) associated with milk production traits.\n\nConclusionsUsing novel analytical approaches, we report the first identification of numerous bovine sQTLs which are extensively shared between multiple tissue types. The significant overlaps between bovine sQTLs and complex traits QTL highlight the contribution of regulatory mutations to phenotypic variations.

genomics

Meta-Analysis Of Sequence-Based Association Studies Across Three Cattle Breeds Reveals 25 QTL For Fat And Protein Percentages In Milk At Nucleotide Resolution

BackgroundGenotyping and whole-genome sequencing data have been collected in many cattle breeds. The compilation of large reference panels facilitates imputing sequence variant genotypes for animals that have been genotyped using dense genotyping arrays. Association studies with imputed sequence variant genotypes allow characterization of quantitative trait loci (QTL) at nucleotide resolution particularly when individuals from several breeds are included in the mapping populations.\n\nResultsWe imputed genotypes for more than 28 million sequence variants in 17,229 animals of the Braunvieh (BV), Fleckvieh (FV) and Holstein (HOL) cattle breeds in order to generate large mapping populations that are required to identify sequence variants underlying milk production traits. Within-breed association tests between imputed sequence variant genotypes and fat and protein percentages in milk uncovered between six and thirteen QTL (P<1e-8) per breed. Eight of the detected QTL were significant in more than one breed. We combined the association studies across three breeds using meta-analysis and identified 25 QTL including six that were not significant in the within-breed association studies. Closer inspection of the QTL revealed that two well-known causal missense mutations in the ABCG2 (p.Y581S, rs43702337, P=4.3e-34) and GHR (p.F279Y, rs385640152, P=1.6e-74) genes were the top variants at two QTL on chromosomes 6 and 20. Another true causal missense mutation in the DGAT1 gene (p.A232K, rs109326954, P=8.4e-1436) was the second top variant at a QTL on chromosome 14 but its allelic substitution effects were not consistent across three breeds analyzed. It turned out that the conflicting allelic substitution effects resulted from flaws in the imputed genotypes due to the use of a multi-breed reference population for genotype imputation.\n\nConclusionsMany QTL for milk production traits segregate across breeds. Metaanalysis of association studies across breeds has greater power to detect such QTL than within-breed association studies. True causal mutations can be readily detected among the most significantly associated variants at QTL when the accuracy of imputation is high. However, true causal mutations may show conflicting allelic substitution effects across breeds when the imputed sequence variant genotypes contain flaws. Validating the effect of known causal variants is highly recommended in order to assess the ability to detect true causal mutations in association studies with imputed sequence variant genotypes.

genomics

Evaluation of the accuracy of imputed sequence variants and their utility for causal variant detection in cattle

BackgroundThe availability of dense genotypes and whole-genome sequence variants from various sources offers the opportunity to compile large data sets consisting of tens of thousands of animals with genotypes for millions of polymorphic sites that may enhance the power of genomic analyses. The imputation of missing genotypes ensures that all individuals have genotypes for a shared set of variants.\n\nResultsWe evaluated the accuracy of imputation from dense genotypes to whole-genome sequence variants in 249 Fleckvieh and 450 Holstein cattle using Minimac and FImpute. The sequence variants of a subset of the animals were reduced to the variants that were included in the Illumina BovineHD genotyping array and subsequently inferred in silico using either within- or multi-breed reference populations. The accuracy of imputation varied considerably across chromosomes and dropped at regions where the bovine genome contains segmental duplications. Depending on the imputation strategy, the correlation between imputed and true genotypes ranged from 0.898 to 0.952. The accuracy of imputation was higher with Minimac than FImpute particularly for rare alleles. Considering a multi-breed reference population increased the accuracy of imputation, particularly when FImpute was used to infer genotypes. When the sequence variants were imputed using Minimac, the true genotypes were more correlated to predicted allele dosages than best-guess genotypes. The computing costs to impute 23,256,743 sequence variants in 6958 animals were 10-fold higher with Minimac than FImpute. Association studies with imputed sequence variants revealed seven quantitative trait loci (QTL) for milk fat percentage. Two known causal mutations in the DGAT1 and GHR genes were the most significantly associated variants at two QTL on chromosomes 14 and 20 when Minimac was used to infer genotypes.\n\nConclusionsThe population-based imputation of millions of sequence variants in large cohorts provides accurate genotypes and is computationally feasible. Using a reference population that includes individuals from many breeds increases the accuracy of imputation particularly at low-frequency variants. Considering allele dosages rather than best-guess genotypes as explanatory variables is advantageous for association studies with imputed sequence variants.

genomics