Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

The Unreasonable Effectiveness of Convolutional Neural Networks in Population Genetic Inference

Population-scale genomic datasets have given researchers incredible amounts of information from which to infer evolutionary histories. Concomitant with this flood of data, theoretical and methodological advances have sought to extract information from genomic sequences to infer demographic events such as population size changes and gene flow among closely related populations/species, construct recombination maps, and uncover loci underlying recent adaptation. To date most methods make use of only one or a few summaries of the input sequences and therefore ignore potentially useful information encoded in the data. The most sophisticated of these approaches involve likelihood calculations, which require theoretical advances for each new problem, and often focus on a single aspect of the data (e.g. only allele frequency information) in the interest of mathematical and computational tractability. Directly interrogating the entirety of the input sequence data in a likelihood-free manner would thus offer a fruitful alternative. Here we accomplish this by representing DNA sequence alignments as images and using a class of deep learning methods called convolutional neural networks (CNNs) to make population genetic inferences from these images. We apply CNNs to a number of evolutionary questions and find that they frequently match or exceed the accuracy of current methods. Importantly, we show that CNNs perform accurate evolutionary model selection and parameter estimation, even on problems that have not received detailed theoretical treatments. Thus, when applied to population genetic alignments, CNN are capable of outperforming expert-derived statistical methods, and offer a new path forward in cases where no likelihood approach exists.

evolutionary biology

A quasi-integral controller for adaptation of genetic modules to variable ribosome demand

The behavior of genetic circuits is often poorly predictable. A genes expression level is not only determined by the intended regulators, but also largely dictated by changes in ribosome availability imparted by activation or repression of other genes. To address this problem, we design a quasi-integral biomolecular feedback controller that enables the expression level of any gene of interest (GOI) to adapt to changes in available ribosomes. The feedback is implemented through a synthetic small RNA (sRNA) that silences the GOIs mRNA, and uses orthogonal extracytoplasmic function (ECF) sigma factor to sense the GOIs translation and to actuate sRNA transcription. Without the controller, the expression level of the GOI is reduced by 50% when a resource competitor is activated. With the controller, by contrast, gene expression level is practically unaffected by the competitor. This feedback controller allows adaptation of genetic modules to variable ribosome demand and thus aids modular construction of complicated circuits.

synthetic biology

MHC GENETIC VARIATION INFLUENCES BOTH OLFACTORY SIGNALS AND SCENT DISCRIMINATION IN RING-TAILED LEMURS

Diversity at the Major Histocompatibility Complex (MHC) is critical to health and fitness, such that MHC genotype may predict an individuals quality or compatibility as a competitor, ally, or mate. Moreover, because MHC products can influence the components of bodily secretions, an individuals body odor may signal its MHC and influence partner identification or mate choice. To investigate MHC-based signaling and recipient sensitivity, we test for odor-gene covariance and behavioral discrimination of MHC diversity and pairwise dissimilarity, under the good genes and good fit paradigms, in a strepsirrhine primate, the ring-tailed lemur (Lemur catta). First, we coupled genotyping with gas chromatography-mass spectrometry to investigate if diversity of the MHC-DRB gene is signaled by the chemical diversity of lemur genital scent gland secretions. We also assessed if the chemical similarity between individuals correlated with their MHC similarity. Next, we assessed if lemurs discriminated this chemically encoded, genetic information in opposite-sex conspecifics. We found that both sexes signaled overall MHC diversity and pairwise MHC similarity via genital secretions, but in a sex- and season-dependent manner. Additionally, both sexes discriminated absolute and relative MHC-DRB diversity in the genital odors of opposite-sex conspecifics, supporting previous findings that lemur genital odors function as advertisement of genetic quality. In this species, genital odors provide honest information about an individuals absolute and relative MHC quality. Complementing evidence in humans and Old World monkeys, our results suggest that reliance on scent signals to communicate MHC quality may be important across the primate lineage.

animal behavior and cognition

Genetic diversity and distribution of indigenous soybean-nodulating bradyrhizobia in the Philippines

The diversity of indigenous bradyrhizobia from soils collected at 11 locations in the Philippines was investigated using PSB-SY2 local soybean cultivar as the host plant. Polymerase Chain Reaction-Restriction Fragment Length Polymorphism (PCR-RFLP) treatment for 16S rRNA, 16S-23S rRNA internal transcribed spacer (ITS) region and rpoB housekeeping gene was performed primarily to detect the genetic variation among the 424 isolates collected. Then, sequence analysis of 16S rRNA, ITS region and rpoB gene was performed for the representative isolates. Majority of the isolates were classified under Bradyrhizobium elkanii, B. diazoefficiens, B. japonicum, Bradyrhizobium sp., and few isolates were related to B. yuanmingense. Genetic variations observed through PCR-RFLP and sequence analyses of the ITS region and rpoB gene generally occurred in B. elkanii, suggesting an occurrence of gene transfer. Shannons diversity index showed varied results with a lowest score of 0.00 and highest at 0.98 indicating a very diverse population of bradyrhizobia across the country. Among all the factors considered in this work, soil management such as period of flooding and some soil properties provided major influence on the distribution and diversity of soybean bradyrhizobia in the country. Thus, it is proposed that the major micro-symbiont of soybean in the Philippines are B. elkanii for non-flooded soils, then B. diazoefficiens and B. japonicum for flooded soils.\n\nImportanceAgriculture production in the Philippines has been and is currently heavily dependent on chemical inputs with mainly rice or corn mono-cropping that it rendered the soil acidic and unproductive. Legume research in the country are mainly focused on plant varietal improvements and very few are aimed at understanding the ecological niche of rhizobia present in the soil. Since soybean has mutual relationship with rhizobia, this legume is a good fallow crop or a rotation crop after rice and corn to help build up the nitrogen stock in the soil. The significance of this research is the better understanding of the ecological niche of indigenous soybean bradyrhizobia, particularly in a tropical archipelago like the Philippines. This work was conceptualized with the utmost goal to increase soybean yield by harnessing and evaluating the indigenous rhizobia in the soil to make production more sustainable and human-friendly.

microbiology

Whole genome sequencing in multiplex families reveals novel inherited and de novo genetic risk in autism

Genetic studies of autism spectrum disorder (ASD) have revealed a complex, heterogeneous architecture, in which the contribution of rare inherited variation remains relatively un-explored. We performed whole-genome sequencing (WGS) in 2,308 individuals from families containing multiple affected children, including analysis of single nucleotide variants (SNV) and structural variants (SV). We identified 16 new ASD-risk genes, including many supported by inherited variation, and provide statistical support for 69 genes in total, including previously implicated genes. These risk genes are enriched in pathways involving negative regulation of synaptic transmission and organelle organization. We identify a significant protein-protein interaction (PPI) network seeded by inherited, predicted damaging variants disrupting highly constrained genes, including members of the BAF complex and established ASD risk genes. Analysis of WGS also identified SVs effecting non-coding regulatory regions in developing human brain, implicating NR3C2 and a recurrent 2.5Kb deletion within the promoter of DLG2. These data lend support to studying multiplex families for identifying inherited risk for ASD. We provide these data through the Hartwell Autism Research and Technology Initiative (iHART), an open access cloud-computing repository for ASD genetics research.

genomics

An Atlas of Human and Murine Genetic Influences on Osteoporosis

Osteoporosis is a common debilitating chronic disease diagnosed primarily using bone mineral density (BMD). We undertook a comprehensive assessment of human genetic determinants of bone density in 426,824 individuals, identifying a total of 518 genome-wide significant loci, (301 novel), explaining 20% of the total variance in BMD--as estimated by heel quantitative ultrasound (eBMD). Next, meta-analysis identified 13 bone fracture loci in ~1.2M individuals, which were also associated with BMD. We then identified target genes from cell-specific genomic landscape features, including chromatin conformation and accessible chromatin sites, that were strongly enriched for genes known to influence bone density and strength (maximum odds ratio = 58, P = 10-75). We next performed rapid throughput skeletal phenotyping of 126 knockout mice lacking eBMD Target Genes and showed that these mice had an increased frequency of abnormal skeletal phenotypes compared to 526 unselected lines (P < 0.0001). In-depth analysis of one such Target Gene, DAAM2, showed a disproportionate decrease in bone strength relative to mineralization. This comprehensive human and murine genetic atlas provides empirical evidence testing how to link associated SNPs to causal genes, offers new insights into osteoporosis pathophysiology and highlights opportunities for drug development.

genomics

Phase-type distributions in population genetics

Probability modelling for DNA sequence evolution is well established and provides a rich framework for understanding genetic variation between samples of individuals from one or more populations. We show that both classical and more recent models for coalescence (with or without recombination) can be described in terms of the so-called phase-type theory, where complicated and tedious calculations are circumvented by the use of matrices. The application of phase-type theory consists of describing the stochastic model as a Markov model by appropriately setting up a state space and calculating the corresponding intensity and reward matrices. Formulae of interest are then expressed in terms of these aforementioned matrices. We illustrate this by a few examples calculating the mean, variance and even higher order moments of the site frequency spectrum in the multiple merger coalescent models, and by analysing the mean and variance for the number of segregating sites for multiple samples in the two-locus ancestral recombination graph. We believe that phase-type theory has great potential as a tool for analysing probability models in population genetics. The compact matrix notation is useful for clarification of current models, in particular their formal manipulation (calculation), but also for further development or extensions.

bioinformatics

Neuropathological correlates and genetic architecture of microglial activation in elderly human brain

Microglia, the resident immune cells of the brain, have important roles in brain health. However, little is known about the regulation and consequences of microglial activation in the aging human brain. We assessed the effect of microglial activation in the aging human brain by calculating the proportion of activated microglia (PAM), based on morphologically defined stages of activation in four regions sampled postmortem from up to 225 elderly individuals. We found that cortical and not subcortical PAM measures were strongly associated with {beta}-amyloid, tau-related neuropathology, and rates of cognitive decline. Effect sizes for PAM measures are substantial, comparable to that of APOE {varepsilon}4, the strongest genetic risk factor for Alzheimers disease. Mediation modeling suggests that PAM accelerates accumulation of tau pathology leading to cognitive decline, supporting an upstream role for microglial activation in Alzheimers disease. Genome-wide analyses identified a common variant (rs2997325) influencing cortical PAM that also affected in vivo microglial activation measured by positron emission tomography using [11C]-PBR28 in an independent cohort. Finally, we identify overlaps of PAMs genetic architecture with those of Alzheimers disease, educational attainment, and several other traits.

genomics

The Genetic Basis for the Cooperative Bioactivation of Plant Lignans by a Human Gut Bacterial Consortium

Plant-derived lignans, consumed daily by most individuals, are inversely associated with breast cancer; however, their bioactivity is only exerted following gut bacterial conversion to enterolignans. Here, we dissect a four-species bacterial consortium sufficient for all four chemical reactions in this pathway. Comparative genomics and heterologous expression experiments identified the first enzyme in the pathway. Transcriptional profiling (RNAseq) independently identified the same gene and linked a single genomic locus to each of the remaining biotransformations. Remarkably, we detected the complete bacterial lignan metabolism pathway in the majority of human gut microbiomes. Together, these results are an important step towards a molecular genetic understanding of the gut bacterial bioactivation of lignans and other plant secondary metabolites to downstream metabolites relevant to human disease.\n\nOne Sentence SummaryBess et al. provide a first step towards elucidating the molecular genetic basis for the cooperative gut bacterial bioactivation of plant lignans, consumed daily by most individuals, to phytoestrogenic enterolignans.

microbiology

Ancestral fitness and genetic background influence diversity acquired during adaptation to low drug concentration in Candida albicans

The importance of within-species diversity in determining the evolutionary potential of a population to evolve drug resistance or tolerance is not well understood, including in eukaryotic pathogens. To examine the influence of genetic background, we evolved replicates of twenty different clinical isolates of Candida albicans, a human fungal pathogen, in fluconazole, the commonly used antifungal drug. The isolates hailed from the major C. albicans clades and had different initial levels of drug resistance and tolerance to the drug. The majority of replicates rapidly increased in fitness in the evolutionary environment, with the degree of improvement inversely correlated with ancestral strain fitness in the drug. Improvement was largely restricted to up to the evolutionary level of drug: only 4% of the evolved replicates increased resistance (MIC) above the evolutionary level of drug. Prevalent changes were altered levels of drug tolerance (slow growth of a subpopulation of cells at drug concentrations above the MIC) and increased diversity of genome size. The prevalence and predominant direction of these changes differed in a strain-specific manner but neither correlated directly with ancestral fitness or improvement in fitness. Rather, low ancestral strain fitness was correlated with high levels of heterogeneity in fitness, tolerance, and genome size among evolved replicates. Thus, ancestral strain background is an important determinant in mean improvement to the evolutionary environment as well as the diversity of evolved phenotypes, and the range of possible responses of a pathogen to an antimicrobial drug cannot be captured by in-depth study of a single strain background. ImportanceAntimicrobial resistance is an evolutionary phenomenon with clinical implications. We tested how replicates from diverse strains of Candida albicans, a prevalent human fungal pathogen, evolve in the commonly-prescribed antifungal drug fluconazole. Replicates on average increased in fitness in the level of drug they were evolved to, with the least fit ancestral strains improving the most. Very few replicates increased resistance above the drug level they were evolved in. Notably, many replicates increased in genome size and changed in drug tolerance (a drug response where a subpopulation of cells grow slowly in high levels of drug) and variability among replicates in fitness, tolerance and genome size was higher in strains that initially were more sensitive to the drug. Genetic background influenced the average degree of adaptation and the evolved variability of many phenotypes, highlighting that different strains from the same species may respond and adapt very differently during adaptation.

evolutionary biology

Illuminating women’s hidden contribution to the foundation of theoretical population genetics

Plentiful evidence shows an historic and continuing gender gap in participation and success in scientific research. However, less attention has been directed at clarifying obscured contributions of women to science. The lack of visible women role models (particularly in computational fields) contributes to a reduced sense of belonging and retention among women. We seek to counteract this cycle by illuminating the contribution of women programmers to the foundation of our own fields--population and evolutionary genetics. We consider past acknowledged programmers (APs), who developed, ran, and sometimes analyzed the results of early computer programs. Due to authorship norms at the time, these programmers were credited in the acknowledgments sections of manuscripts, rather than being recognized as authors. For example, one acknowledgement reads \"I thanks Mrs. M. Wu for help with the numerical work, and in particular for computing table I.\". We identified APs in Theoretical Population Biology articles published between 1970 and 1990. While only 7% of authors were women, 43% of APs were women. This significant difference (p = 4.0x10-10) demonstrates a substantial proportion of womens contribution to foundational computational population genetics has been unrecognized. The proportion of women APs, as well as number of APs decreased over time. These observations correspond to the masculinization of computer programming, and the shifting of programming responsibilities to individuals credited as authors (likely graduate students). Finally, we note recurrent APs who contributed to several highly-cited manuscripts. We conclude that, while previously overlooked, historically, women have made substantial contributions to computational biology.

scientific communication and education

ClinPhen extracts and prioritizes patientphenotypes directly from medical records to accelerate genetic disease diagnosis

PurposeSevere genetic diseases affect 7 million births per year, worldwide. Diagnosing these diseases is necessary for optimal care, but it can involve the manual evaluation of hundreds of genetic variants per case, with many variants taking an hour to evaluate. Automatic gene-ranking approaches shorten this process by reporting which of the genes containing variants are most likely to be causing the patients symptoms. To use these tools, busy clinicians must manually encode patient phenotypes, which is a cumbersome and imprecise process. With 60 million patients expected to be sequenced in the next 7 years, a fast alternative to manual phenotype extraction from the clinical notes in patients medical records will become necessary.\n\nMethodsWe introduce ClinPhen: a fast, high-accuracy tool that automatically converts the clinical notes into a prioritized list of patient symptoms using HPO terms.\n\nResultsClinPhen shows superior accuracy to existing phenotype extractors, and when paired with a gene-ranking tool it significantly improve the latters performance.\n\nConclusionCompared to manual phenotype extraction, ClinPhen saves more than 5 hours per case in Mendelian diagnosis alone. Summing over millions of forthcoming cases whose medical notes await phenotype encoding, ClinPhen makes a substantial contribution towards ending all patients diagnostic odyssey.

genomics

A non-coding genetic variant maximally associated with serum urate levels is functionally linked to HNF4A-dependent PDZK1 expression

Several dozen genetic variants associate with serum urate levels, but the precise molecular mechanisms by which they affect serum urate are unknown. Here we tested for functional linkage of the maximally-associated genetic variant rs1967017 at the PDZK1 locus to elevated PDZK1 expression.\n\nWe performed expression quantitative trait locus (eQTL) and likelihood analyses followed by gene expression assays. Zebrafish were used to determine the ability of rs1967017 to direct tissue-specific gene expression. Luciferase assays in HEK293 and HepG2 cells measured the effect of rs1967017 on transcription amplitude.\n\nPAINTOR analysis revealed rs1967017 as most likely to be causal and rs1967017 was an eQTL for PDZK1 in the intestine. The region harboring rs1967017 was capable of directly driving green fluorescent protein expression in the kidney, liver and intestine of zebrafish embryos, consistent with a conserved ability to confer tissue-specific expression. The urate-increasing T-allele of rs1967017 strengthens a binding site for the transcription factor HNF4A. siRNA depletion of HNF4A reduced endogenous PDZK1 expression in HepG2 cells. Luciferase assays showed that the T-allele of rs1967017 gains enhancer activity relative to the urate-decreasing C-allele, with T-allele enhancer activity abrogated by HNF4A depletion. HNF4A physically binds the rs1967017 region, suggesting direct transcriptional regulation of PDZK1 by HNF4A.\n\nWith other reports our data predict that the urate-raising T-allele of rs1967017 enhances HNF4A binding to the PDZK1 promoter, thereby increasing PDZK1 expression. As PDZK1 is a scaffold protein for many ion channel transporters, increased expression can be predicted to increase activity of urate transporters and alter excretion of urate.

cell biology

Genetic mapping in Diversity Outbred mice identifies a Trpa1 variant influencing late phase formalin response

Identification of genetic variants that influence susceptibility to chronic pain is key to identifying molecular mechanisms and targets for effective and safe therapeutic alternatives to opioids. To identify genes and variants associated with chronic pain, we measured late phase response to formalin injection in 275 male and female Diversity Outbred (DO) mice genotyped for over 70 thousand SNPs. One quantitative trait locus (QTL) reached genome-wide significance on chromosome 1 with a support interval of 3.1 Mb. This locus, Nociq4 (nociceptive sensitivity inflammatory QTL 4; MGI:5661503), harbors the well-known pain gene Trpa1 (transient receptor potential cation channel, subfamily A, member 1). Trpa1 is a cation channel known to play an important role in acute and chronic pain in both humans and mice. Analysis of DO founder strain allele effects revealed a significant effect of the CAST/EiJ allele at Trpa1, with CAST/EiJ carrier mice showing an early, but not late, response to formalin relative to carriers of the seven other inbred founder alleles (A/J, C57BL/6J, 129S1/SvImJ, NOD/ShiLtJ, NZO/HlLtJ, PWK/PhJ, and WSB/EiJ). We characterized possible functional consequences of sequence variants in Trpa1 by assessing channel conductance, Trpa1/Trpv1 interactions, and isoform expression. The phenotypic differences observed in CAST/EiJ relative to C57BL/6J carriers were best explained by Trpa1 isoform expression differences, implicating a splice junction variant as the causal functional variant. This study demonstrates the utility of advanced, high-precision genetic mapping populations in resolving specific molecular mechanisms of variation in pain sensitivity.

bioinformatics

Analysis of head and neck carcinoma progression reveals novel and relevant stage-specific genetic changes associated with immortalisation and malignancy.

Head and neck squamous cell carcinoma (HNSCC) is a widely prevalent cancer globally with high mortality and morbidity. We report here changes in the genomic landscape in the development of HNSCC from potentially premalignant lesions (PPOLS) to malignancy and lymph node metastases. Frequent likely pathological mutations are restricted to a relatively small set of genes including TP53, CDKN2A, FBXW7, FAT1, NOTCH1 and KMT2D; these arise early in tumour progression and are present in PPOLs with NOTCH1 mutations restricted to cell lines from lesions that subsequently progressed to HNSCC. The most frequent genetic changes are of consistent somatic copy number alterations (SCNA). The earliest SCNAs involved deletions of CSMD1 (8p23.2), FHIT (3p14.2) and CDKN2A (9p21.3) together with gains of chromosome 20. CSMD1 deletions or promoter hypermethylation were present in all of the immortal PPOLs and occurred at high frequency in the immortal HNSCC cell lines (promoter hypermethylation ~63%, hemizygous deletions ~75%, homozygous deletions ~18%). Forced expression of CSMD1 in the HNSCC cell line H103 showed significant suppression of proliferation (p=0.0053) and invasion in vitro (p=5.98X10-5) supporting a role for CSMD1 inactivation in early head and neck carcinogenesis. In addition, knockdown of CSMD1 in the CSMD1-expressing BICR16 cell line showed significant stimulation of invasion in vitro (p=1.82 x 10-5) but not cell proliferation (p=0.239). HNSCC with and without nodal metastases showed some clear differences including high copy number gains of CCND1, hsa-miR-548k and TP63 in the metastases group. GISTIC peak SCNA regions showed significant enrichment (adj P<0.01) of genes in multiple KEGG cancer pathways at all stages with disruption of an increasing number of these involved in the progression to lymph node metastases. Sixty-seven genes from regions with statistically significant differences in SCNA/LOH frequency between immortal PPOL and HNSCC cell lines showed correlation with expression including 5 known cancer drivers.\n\nLay SummaryCancers affecting the head and neck region are relatively common. A large percentage of these are of one particular type; these are generally detected late and are associated with poor prognosis. Early detection and treatment dramatically improve survival and reduces the damage associated with the cancer and its treatment. Cancers arise and progress because of changes in the genetic material of the cells. This study focused on identifying such changes in these cancers particularly in the early stages of development, which are not fully known. Identification of these changes is important in developing new treatments as well as markers of behaviour of cancers and also the early or premalignant lesions. We used a well-characterised panel of cell lines generated from premalignant lesions as well as cancers, to identify mutations in genes, and an increase or decrease in number of copies of genes. We mapped new and previously identified changes in these cancers to specific stages in the development of these cancers and their spread. We additionally report here for the first time, alterations in CSMD1 gene in early premalignant lesions; we further show that this is likely to result in increased ability of the cells to spread and possibly, multiply faster as well.

cancer biology

The first genetic linkage map for Fraxinus pennsylvanica and syntenic relationships with four related species

Green ash (Fraxinus pennsylvanica) is an outcrossing, diploid (2n=46) hardwood tree species, native to North America. Native ash species in North America are being threatened by the rapid invasion of emerald ash borer (EAB, Agrilus planipennis) from Asia. Green ash, the most widely distributed ash species, is severely affected by EAB infestation, yet few resources for genetic studies and improvement of green ash are available. In this study, a total of 5,712 high quality single nucleotide polymorphisms (SNPs) were discovered using a minimum allele frequency of 1% across the entire genome through genotyping-by-sequencing. We also screened hundreds of genomic- and EST-based microsatellite markers (SSRs) from previous de novo assemblies (Staton et al. 2015; Lane et al. 2016). A first genetic linkage map of green ash was constructed from 91 individuals in a full-sib family, combining 2,719 SNP and 84 SSR segregating markers among the parental maps. The consensus SNP and SSR map contains a total of 1,201 markers in 23 linkage groups spanning 2008.87cM, at an average inter-marker distance of 1.67 cM with a minimum logarithm of odds (LOD) of 6 and maximum recombination fraction of 0.40. Comparisons of the organization the green ash map with the genomes of asterid species coffee and tomato, and genomes of the rosid species poplar and peach, showed areas of conserved gene order, with overall synteny strongest with coffee.

plant biology

Genetic and environmental perturbations lead to regulatory decoherence

Correlation among traits is a fundamental feature of biological systems. From morphological characters, to transcriptional or metabolic networks, the correlations we routinely observe between traits reflect a shared regulation that remains poorly understood and difficult to study. To address this problem, we developed a new and flexible approach that allows us to identify factors associated with variation in correlation between individuals. Here, we use data from three large human cohorts to study the effects of genetic variation and environmental perturbation on correlations among mRNA transcripts and among NMR metabolites. We first show that environmental exposures (namely, infection and disease) lead to a systematic loss of correlation, which we define as decoherence. Using longitudinal data, we show that decoherent metabolites are better predictors of whether someone will develop metabolic syndrome than metabolites commonly used as biomarkers of this disease. Finally, we show that correlation itself is a trait under genetic control: specifically, we mapped and replicated hundreds of correlation QTLs, which often involve transcription factors or their known target genes. Together, this work furthers our understanding of how and why coordinated biological processes break down, and highlights the role of decoherence in disease emergence.

genomics

Exploring the Genetic Basis of Human Population Differences in DNA Methylation and their Causal Impact on Immune Gene Regulation

DNA methylation is influenced by both environmental and genetic factors and is increasingly thought to affect variation in complex traits and diseases. Yet, the extent of ancestry-related differences in DNA methylation, its genetic determinants, and their respective causal impact on immune gene regulation remain elusive. We report extensive population differences in DNA methylation between individuals of African and European descent -- detected in primary monocytes that were used as a model of a major innate immunity cell type. Most of these differences (~70%) were driven by DNA sequence variants nearby CpG sites (meQTLs), which account for ~60% of the variance in DNA methylation. We also identify several master regulators of DNA methylation variation in trans, including a regulatory hub nearby the transcription factor-encoding CTCF gene, which contributes markedly to ancestry-related differences in DNA methylation. Furthermore, we establish that variation in DNA methylation is associated with varying gene expression levels following mostly, but not exclusively, a canonical model of negative associations, particularly in enhancer regions. Specifically, we find that DNA methylation highly correlates with transcriptional activity of 811 and 230 genes, at the basal state and upon immune stimulation, respectively. Finally, using a Bayesian approach, we estimate causal mediation effects of DNA methylation on gene expression in ~20% of the studied cases, indicating that DNA methylation can play an active role in immune gene regulation. Using a system-level approach, our study reveals substantial ancestry-related differences in DNA methylation and provides evidence for their causal impact on immune gene regulation.

genomics