Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,639 records · Page 91Linked to original sources

HaploVectors: an integrative analytical tool for phylogeography

Sets of local populations show different degrees of gene flow due to dispersal barriers and environmental constraints, which renders genetic composition gradients among populations (genetic turnover). Unveiling biogeographic correlates of genetic turnover is paramount for phylogeography. While some processes (genetic drift, secondary contact) may erase the historical track of genetic turnover, vicariance or ancient dispersal likely leads to genetic divergence among populations. Yet available analyses do not permit direct inference about mechanisms driving genetic turnover. We propose a novel analytical approach called genvector analysis, which fulfills this gap by decomposing genetic compositional dissimilarities between populations based on either haplotypes or other genetic data into genetic eigenvectors. Such procedure allows exploring genetic turnover among sets of local populations, and analyzing their biogeographic correlates based on null model tests. We evaluate the statistical performance of the method on simulated datasets. We also analyzed biogeographic correlates of genetic turnover of Akodon cursor in the Brazilian Atlantic Forest. Results revealed that genvector analysis is robust to discriminate biogeographic drivers of genetic turnover. For Akodon cursor analysis, we observed that while for the entire species, all predictors considered (except for elevation) explained genetic turnover, within phylogroups some factors varied their importance. Genvector analysis was demonstrated to be useful for several different purposes in phylogeography, and complementary to classic analytical tools widely used by phylogeographers, such as AMOVA or DAPC. The role of ancient versus recent biogeographic events, the relationship between morphological divergence or abiotic variables and genetic turnover can easily be investigated using genvector analysis.

ecology

Insular Celtic population structure and genomic footprints of migration

Previous studies of the genetic landscape of Ireland have suggested homogeneity, with population substructure undetectable using single-marker methods. Here we have harnessed the haplotype-based method fineSTRUCTURE in an Irish genome-wide SNP dataset, identifying 23 discrete genetic clusters which segregate with geographical provenance. Cluster diversity is pronounced in the west of Ireland but reduced in the east where older structure has been eroded by historical migrations. Accordingly, when populations from the neighbouring island of Britain are included, a west-east cline of Celtic-British ancestry is revealed along with a particularly striking correlation between haplotypes and geography across both islands. A strong relationship is revealed between subsets of Northern Irish and Scottish populations, where discordant genetic and geographic affinities reflect major migrations in recent centuries. Additionally, Irish genetic proximity of all Scottish samples likely reflects older strata of communication across the narrowest inter-island crossing. Using GLOBETROTTER we detected Irish admixture signals from Britain and Europe and estimated dates for events consistent with the historical migrations of the Norse-Vikings, the Anglo-Normans and the British Plantations. The influence of the former is greater than previously estimated from Y chromosome haplotypes. In all, we paint a new picture of the genetic landscape of Ireland, revealing structure which should be considered in the design of studies examining rare genetic variation and its association with traits.\n\nAuthor summaryA recent genetic study of the UK (People of the British Isles; PoBI) expanded our understanding of population history of the islands, using newly-developed, powerful techniques that harness the rich information embedded in chunks of genetic code called haplotypes. These methods revealed subtle regional diversity across the UK, and, using genetic data alone, timed key migration events into southeast England and Orkney. We have extended these methods to Ireland, identifying regional differences in genetics across the island that adhere to geography at a resolution not previously reported. Our study reveals relative western diversity and eastern homogeneity in Ireland owing to a history of settlement concentrated on the east coast and longstanding Celtic diversity in the west. We show that Irish Celtic diversity enriches the findings of PoBI; haplotypes mirror geography across Britain and Ireland, with relic Celtic populations contributing greatly to haplotypic diversity. Finally, we used genetic information to date migrations into Ireland from Europe and Britain consistent with historical records of Viking and Norman invasions, demonstrating the signatures of these migrations the on modern Irish genome. Our findings demonstrate that genetic structure exists in even small isolated populations, which has important implications for population-based genetic association studies.

genetics

Evidence of causal effect of major depression on alcohol dependence: Findings from the Psychiatric Genomics Consortium

BackgroundDespite established clinical associations among major depression (MD), alcohol dependence (AD), and alcohol consumption (AC), the nature of the causal relationship between them is not completely understood.\n\nMethodsThis study was conducted using genome-wide data from the Psychiatric Genomics Consortium (MD: 135,458 cases and 344,901 controls; AD: 10,206 cases and 28,480 controls) and UK Biobank (AC-Frequency: from \"daily or almost daily\" to \"never\", 438,308 individuals; AC-Quantity: total units of alcohol per week, 307,098 individuals). Linkage disequilibrium score regression and Mendelian Randomization (MR) analyses were applied to investigate shared genetic mechanisms (horizontal pleiotropy) and causal relationships (mediated pleiotropy) among these traits.\n\nOutcomesPositive genetic correlation was observed between MD and AD (rgMD-AD=+0.47, P=6.6x10-10). AC-Quantity showed positive genetic correlation with both AD (rgAD-AC-Quantity=+0.75, P=1.8x10-14) and MD (rgMD-AC-Quantity=+0.14, P=2.9x10-7), while there was negative correlation of AC-Frequency with MD (rgMD-AC-Frequency=-0.17, P=1.5x10-10) and a non-significant result with AD. MR analyses confirmed the presence of pleiotropy among these traits. However, the MD-AD results reflect a mediated-pleiotropy mechanism (i.e., causal relationship) with a causal role of MD on AD (beta=0.28, P=1.29x10-6) that does not appear to be biased by confounding such as horizontal pleiotropy. No evidence of reverse causation was observed as the AD genetic instrument did not show a causal effect on MD.\n\nInterpretationResults support a causal role for MD on AD based on genetic datasets including thousands of individuals. Understanding mechanisms underlying MD-AD comorbidity not only addresses important public health concerns but also has the potential to facilitate prevention and intervention efforts.\n\nFundingNational Institute of Mental Health and National Institute on Drug Abuse.\n\nPutting data into contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed up to August 24, 2018, for research studies that investigated causality among alcohol-and depression related phenotypes using Mendelian randomization approaches. We used the search terms \"alcohol\" AND \"depression\" AND \"Mendelian Randomization\". No restrictions were applied to language, date, or article type. Ten articles were retrieved, but only two were focused on alcohol consumption and depression-related traits. The studies were based on genetic variants in alcohol dehydrogenase (ADH) genes only, did not find evidence for a causal effect of alcohol consumption on depression phenotypes, with one study finding a causal effect of alcohol consumption on alcoholism. Both studies noted that future studies are needed with increased sample sizes and clinically derived phenotypes. To our knowledge, no previous study has applied two-sample Mendelian randomization to investigate causal relationships between alcohol dependence and major depression.\n\nTwin studies show genetic factors influence susceptibility to MD, AD, and alcohol consumption. Differently from observational approaches where several studies have investigated the relationship between alcohol-and depression-related phenotypes, very limited use of molecular genetic data has been applied to investigate this issue. Additionally, the use of genetic information has been shown to be less biased by confounders and reverse causation than observation data. However, genetic approaches, like Mendelian randomization, require large sample sizes to be informative.\n\nAdded value of this studyIn this study, we used genome-wide data from the Psychiatric Genomic Consortium and UK Biobank, which include information regarding hundred thousands of individuals, to test the presence of shared genetic mechanisms and causal relationships among major depression, alcohol dependence, and alcohol consumption. The results support a causal influence of MD on AD, while alcohol consumption showed shared genetic mechanisms with respect to both major depression and alcohol dependence.\n\nImplications of all the available evidenceGiven the significant morbidity and mortality associated with MD, AD, and the comorbid condition, understanding mechanisms underlying these associations not only address important public health concerns but also has the potential to facilitate prevention and intervention efforts.

genetics

Candidate gene scan for Single Nucleotide Polymorphisms involved in the determination of normal variability in human craniofacial morphology

Despite intensive research on genetics of the craniofacial morphology using animal models and human craniofacial syndromes, the genetic variation that underpins normal human facial appearance is still largely elusive. Recent development of novel digital methods for capturing the complexity of craniofacial morphology in conjunction with high-throughput genotyping methods, show great promise for unravelling the genetic basis of such a complex trait.\n\nAs a part of our efforts on detecting genomic variants affecting normal craniofacial appearance, we have implemented a candidate gene approach by selecting 1,201 single nucleotide polymorphisms (SNPs) and 4,732 tag SNPs in over 170 candidate genes and intergenic regions. We used 3-dimentional (3D) facial scans and direct cranial measurements of 587 volunteers to calculate 104 craniofacial phenotypes. Following genotyping by massively parallel sequencing, genetic associations between 2,332 genetic markers and 104 craniofacial phenotypes were tested.\n\nAn application of a Bonferroni-corrected genome-wide significance threshold produced significant associations between five craniofacial traits and six SNPs. Specifically, associations of nasal width with rs8035124 (15q26.1), cephalic index with rs16830498 (2q23.3), nasal index with rs37369 (5q13.2), transverse nasal prominence angle with rs59037879 (10p11.23) and rs10512572 (17q24.3), and principal component explaining 73.3% of all the craniofacial phenotypes, with rs37369 (5p13.2) and rs390345 (14q31.3) were observed.\n\nDue to over-conservative nature of the Bonferroni correction, we also report all the associations that reached the traditional genome-wide p-value threshold (<5.00E-08) as suggestive. Based on the genome-wide threshold, 8 craniofacial phenotypes demonstrated significant associations with 34 intergenic and extragenic SNPs. The majority of associations are novel, except PAX3 and COL11A1 genes, which were previously reported to affect normal craniofacial variation.\n\nThis study identified the largest number of genetic variants associated with normal variation of craniofacial morphology to date by using a candidate gene approach, including confirmation of the two previously reported genes. These results enhance our understanding of the genetics that determines normal variation in craniofacial morphology and will be of particular value in medical and forensic fields.\n\nAuthor SummaryThere is a remarkable variety of human facial appearances, almost exclusively the result of genetic differences, as exemplified by the striking resemblance of identical twins. However, the genes and specific genetic variants that affect the size and shape of the cranium and the soft facial tissue features are largely unknown. Numerous studies on animal models and human craniofacial disorders have identified a large number of genes, which may regulate normal craniofacial embryonic development.\n\nIn this study we implemented a targeted candidate gene approach to select more than 1,200 polymorphisms in over 170 genes that are likely to be involved in craniofacial development and morphology. These markers were genotyped in 587 DNA samples using massively parallel sequencing and analysed for association with 104 traits generated from 3-dimensional facial images and direct craniofacial measurements. Genetic associations (p-values<5.00E-08) were observed between 8 craniofacial traits and 34 single nucleotide polymorphisms (SNPs), including two previously described genes and 26 novel candidate genes and intergenic regions. This comprehensive candidate gene study has uncovered the largest number of novel genetic variants affecting normal facial appearance to date. These results will appreciably extend our understanding of the normal and abnormal embryonic development and impact our ability to predict the appearance of an individual from a DNA sample in forensic criminal investigations and missing person cases.

Genetics

Human demographic history has amplified the effects background selection across the genome

Natural populations often grow, shrink, and migrate over time. Demographic processes such as these can impact genome-wide levels of genetic diversity. In addition, genetic variation in functional regions of the genome can be altered by natural selection, which drives adaptive mutations to higher frequencies or purges deleterious ones. Such selective processes impact not only the sites directly under selection but also nearby neutral variation through genetic linkage through processes referred to as genetic hitch-hiking in the context of positive selection and background selection (BGS) in the context of purifying selection. While there is extensive literature examining the impact of selection at linked sites at demographic equilibrium, less is known about how non-equilibrium demographic processes impact the effects of hitchhiking and BGS. Utilizing a global sample of human whole-genome sequences from the Thousand Genomes Project and extensive simulations, we investigate how non-equilibrium demographic processes magnify and dampen the consequences of selection at linked sites across the human genome. When binning the genome by inferred strength of BGS, we observe that, compared to Africans, non-African populations have experienced larger proportional decreases in neutral genetic diversity in such regions. We replicate these findings in admixed populations by showing that non-African ancestral components of the genome have also been impacted more severely in these regions. We attribute these differences to the strong, sustained/recurrent population bottlenecks that non-Africans experienced as they migrated out of Africa and throughout the globe. Furthermore, we observe a strong correlation between FST and inferred strength of BGS, suggesting a stronger rate of genetic drift. Forward simulations of human demographic history with a model of BGS support these observations. Our results show that non-equilibrium demography significantly alters the consequences selection at linked sites and support the need for more work investigating the dynamic process of multiple evolutionary forces operating in concert.\n\nAuthor summaryPatterns of genetic diversity within a species are affected at broad and fine scales by population size changes (\"demography\") and natural selection. From both population genetics theory and observation of genomic sequence data, it is known that demography can alter genome-wide average neutral genetic diversity. Additionally, natural selection can affect neutral genetic diversity regionally across the genome via selection at linked sites. During this process, natural selection acting on adaptive or deleterious variants in the genome will also impact diversity at nearby neutral sites due to genetic linkage. However, less is well known about the dynamic changes to diversity that occur in regions impacted by selection at linked sites when a population undergoes a size change. We characterize these dynamic changes using thousands of human genomes and find that the population size changes experienced by humans have shaped the consequences of linked selection across the genome. In particular, population contractions, such as those experienced by non-Africans, have disproportionately decreased neutral diversity in regions of the genome inferred to be under strong background selection (i.e., selection at linked sites that is caused by natural selection acting on deleterious variants), resulting in large differences between African and non-African populations.

evolutionary biology

Neurobehavioural Correlates of Obesity are Largely Heritable

Recent molecular genetic studies have shown that the majority of genes associated with obesity are expressed in the central nervous system. Obesity has also been associated with neurobehavioural factors such as brain morphology, cognitive performance, and personality. Here, we tested whether these neurobehavioural factors were associated with the heritable variance in obesity measured by body mass index (BMI) in the Human Connectome Project (N=895 siblings). Phenotypically, cortical thickness findings supported the \"right brain hypothesis\" for obesity. Namely, increased BMI associated with decreased cortical thickness in right frontal lobe and increased thickness in the left frontal lobe, notably in lateral prefrontal cortex. In addition, lower thickness and volume in entorhinal-parahippocampal structures, and increased thickness in parietal-occipital structures in obese participants supported the role of visuospatial function in obesity. Brain morphometry results were supported by cognitive tests, which outlined obesitys negative association with visuospatial function, verbal episodic memory, impulsivity, and cognitive flexibility. Personality-obesity correlations were inconsistent. We then aggregated the effects for each neurobehavioural factor for a behavioural genetics analysis and demonstrated the factors genetic overlap with obesity. Namely, cognitive test scores and brain morphometry had 0.25 - 0.45 genetic correlations with obesity, and the phenotypic correlations with obesity were 77-89% explained by genetic factors. Neurobehavioural factors also had some genetic overlap with each other. In summary, obesity has considerable genetic overlap with brain and cognitive measures. This supports the theory that obesity is inherited via brain function, and may inform intervention strategies.\n\nSignificance StatementObesity is a widespread heritable health condition. Evidence from psychology, cognitive neuroscience, and genetics has proposed links between obesity and the brain. The current study tested whether the heritable variance in obesity is explained by brain and behavioural factors in a large brain imaging cohort that included multiple related individuals. We found that the heritable variance in obesity had genetic correlations 0.25 - 0.45 with cognitive tests, cortical thickness, and regional brain volume. In particular, obesity was associated with frontal lobe asymmetry and differences in temporal-parietal perceptual systems. Further, we found genetic overlap between certain brain and behavioural factors. In summary, the genetic vulnerability to obesity is expressed in the brain. This may inform intervention strategies.

neuroscience

Aligning functional network constraint to evolutionary outcomes

It is likely that there are constraints on how evolution can progress, and well-known evolutionary phenomena such as convergent evolution, rapid adaptation, and genic evolution would be difficult to explain under the absence of any such evolutionary constraint. One dimension of constraint results from a finite number of environmental conditions, and thus natural selection scenarios, leading to convergent phenotypes. This limits which genetic variants are adaptive, and consequently, constrains how variation is inherited across generations. Another, less explored dimension of evolution is functional constraint at the molecular level. Some widely accepted examples for this dimension of evolutionary constraint include genetic linkage, codon position, and architecture of developmental genetic pathways, that together constrain how evolution can shape genomes through limiting which mutations can increase fitness. Genomic architecture, which describes how all gene products interact, has been discussed to be another dimension of functional genetic constraint. This notion had been largely discredited by the modern synthesis, especially because macroevolution was not always found to be perfectly deterministic. But debates on whether evolutionary constraint stems mostly from environmental (extrinsic) or genetic (intrinsic) factors have mostly been held at the intellectual level using sporadic evidence. Quantifying the relative contributions of these different dimensions of constraint is, however, fundamentally important to understand the mechanistic basis of seemingly deterministic evolutionary outcomes. In some model organisms, genetic constraint has already been quantitatively explored. Forays into testing the relationship between genomic architecture and evolution included studies on protein evolutionary rate variation in essential versus nonessential genes, and observations that the number of protein interactions within a cell (gene pleiotropy) determines the fitness effect of mutations. In this contribution, existing evidence for functional genetic constraint as shaping evolutionary outcomes is reviewed and testable hypotheses are defined for functional genetic constraint influencing (i) convergent evolution, (ii) rapid adaptation, and genic adaptation. An analysis of the yeast interactome incorporating recently published data on its evolution, reveals new support for the existence of genomic architecture as a functional genetic dimension of evolutionary constraint. As functional genetic networks are becoming increasingly available, evolutionary biologists should strive to evaluate functional genetic network constraint, against variables describing complex phenotypes and environments, for better understanding commonly observed deterministic patterns of evolution in non-model organisms. This may help to quantify the extrinsic versus intrinsic dimensions of evolutionary constraint, and result in a better understanding of how fast, effectively, or deterministically organisms adapt.\n\nGlossary

evolutionary biology

Visualizing spatial population structure with estimated effective migration surfaces

Genetic data often exhibit patterns that are broadly consistent with \"isolation by distance\" - a phenomenon where genetic similarity tends to decay with geographic distance. In a heterogeneous habitat, decay may occur more quickly in some regions than others: for example, barriers to gene flow can accelerate the genetic differentiation between groups located close in space. We use the concept of \"effective migration\" to model the relationship between genetics and geography: in this paradigm, effective migration is low in regions where genetic similarity decays quickly. We present a method to quantify and visualize variation in effective migration across the habitat, which can be used to identify potential barriers to gene flow, from geographically indexed large-scale genetic data. Our approach uses a population genetic model to relate underlying migration rates to expected pairwise genetic dissimilarities, and estimates migration rates by matching these expectations to the observed dissimilarities. We illustrate the potential and limitations of our method using simulations and geo-referenced genetic data from elephant, human and Arabidopsis thaliana populations. The resulting visualizations highlight important features of the spatial population structure that are difficult to discern using existing methods for summarizing genetic variation such as principal components analysis.

Genetics

Simple multi-trait analysis identifies novel loci associated with growth and obesity measures

The ever-growing genome-wide association studies (GWAS) have revealed widespread pleiotropy. To exploit this, various methods which consider variant association with multiple traits jointly have been developed. However, most effort has been put on improving discovery power: how to replicate and interpret these discovered pleiotropic loci using multivariate methods has yet to be discussed fully. Using only multiple publicly available single-trait GWAS summary statistics, we develop a fast and flexible multi-trait framework that contains modules for (i) multi-trait genetic discovery, (ii) replication of locus pleiotropic profile, and (iii) multi-trait conditional analysis. The procedure is able to handle any level of sample overlap. As an empirical example, we discovered and replicated 23 novel pleiotropic loci for human anthropometry and evaluated their pleiotropic effects on other traits. By applying conditional multivariate analysis on the 23 loci, we discovered and replicated two additional multi-trait associated SNPs. Our results provide empirical evidence that multi-trait analysis allows detection of additional, replicable, highly pleiotropic genetic associations without genotyping additional individuals. The methods are implemented in a free and open source R package MultiABEL.\n\nAuthor summaryBy analyzing large-scale genomic data, geneticists have revealed widespread pleiotropy, i.e. single genetic variation can affect a wide range of complex traits. Methods have been developed to discover such genetic variants. However, we still lack insights into the relevant genetic architecture - What more can we learn from knowing the effects of these genetic variants?\n\nHere, we develop a fast and flexible statistical analysis procedure that includes discovery, replication, and interpretation of pleiotropic effects. The whole analysis pipeline only requires established genetic association study results. We also provide the mathematical theory behind the pleiotropic genetic effects testing.\n\nMost importantly, we show how a replication study can be essential to reveal new biology rather than solely increasing sample size in current genomic studies. For instance, we show that, using our proposed replication strategy, we can detect the difference in genetic effects between studies of different geographical origins.\n\nWe applied the method to the GIANT consortium anthropometric traits to discover new genetic associations, replicated in the UK Biobank, and provided important new insights into growth and obesity.\n\nOur pipeline is implemented in an open-source R package MultiABEL, sufficiently efficient that allows researchers to immediately apply on personal computers in minutes.

Genetics

Analytical and Clinical Validity Study of FirstStepDx PLUS: A Chromosomal Microarray Optimized for Patients with Neurodevelopmental Conditions

IntroductionChromosomal microarray analysis (CMA) is recognized as the first-tier test in the genetic evaluation of children with developmental delays, intellectual disabilities, congenital anomalies and autism spectrum disorders of unknown etiology.\n\nArray DesignTo optimize detection of clinically relevant copy number variants associated with these conditions, we designed a whole-genome microarray, FirstStepDx PLUS (FSDX). A set of 88,435 custom probes was added to the Affymetrix CytoScanHD platform targeting genomic regions strongly associated with these conditions. This combination of 2,784,985 total probes results in the highest probe coverage and clinical yield for these disorders.\n\nResults and DiscussionClinical testing of this patient population is validated on DNA from either non-invasive buccal swabs or traditional blood samples. In this report we provide data demonstrating the analytic and clinical validity of FSDX and provide an overview of results from the first 7,570 consecutive patients tested clinically. We further demonstrate that buccal sampling is an effective method of obtaining DNA samples, which may provide improved results compared to traditional blood sampling for patients with neurodevelopmental disorders who exhibit somatic mosaicism.\n\nClinical scenarioNeurodevelopmental disabilities, including developmental delays (DD), intellectual disabilities (ID), and autism spectrum disorders (ASD), affect up to 15% of children (1). In the majority of cases, a childs clinical presentation does not allow for a definitive etiological diagnosis. In such cases, CMA is recommended as the first-tier test that should be used to evaluate for a potential genetic etiology (2-7). A definitive genetic diagnosis allows patients to more often receive appropriate medical care tailored to their condition, as reflected by medical management changes and improved access to necessary support and educational services (8-13).\n\nTest descriptionFirstStepDx PLUS (FSDX) is an optimized clinical microarray test provided in the context of a comprehensive clinical service. Testing starts with either a non-invasive buccal swab sample or traditional blood sample from which DNA extraction using a Gentra Puregene(R) kit specific to the sample type (Qiagen, Inc., Valencia, CA) is performed in one of several contracted CLIA/CAP credentialed laboratories according to manufacturers protocols. High quality genomic DNA is fragmented, labeled and hybridized to FSDX arrays using reagents, equipment and methodology as specified by by the manufacturer (Affymetrix, Inc., Santa Clara, CA)(14). Washed arrays are scanned and raw data files are processed to CYCHP files using a reference file comprising at least 100 samples with normal array findings. Data analysis is performed using Chromosome Analysis Suite software version 2.0.1 (Affymetrix). Hybridization of patient DNA to oligonucleotide and SNP probes is independently compared against a previously analyzed cohort of normal samples to call CNVs and allele genotypes. The percentage mosaicism of whole-chromosome aneuploidies is determined using the average log2 ratio of the entire chromosome (14).\n\nMicroarray designFSDX was optimized by the addition of 88,435 custom probes targeting genomic regions strongly associated with ID/DD/ASD (15-24). This was effected, under GMP by Affymetrix, to the CytoScanHD platform using their microarray design process specifications which have been previously described (14). This is consistent with the ACMG recommendation of \"enrichment of probes targeting dosage-sensitive genes known to result in phenotypes consistent with common indications for a genomic screen\" (25). Critical regions that did not meet a desired probe density [&ge;]1 probe/1000 bp on the CytoScanHD were supplemented with additional probe content to allow for improved detection of smaller deletions and duplications in these critical regions. Finally, additional probes were added to improve detection of CNVs encompassing genes involved in other well-characterized neurodevelopmental disorders, for example GAMT (26) and GATM (27). All incremental probes were added in substitution for probes deemed sub-optimal by Affymetrix and previously masked, bringing FSDX to a grand total of 2,784,985 probes. Custom SNP probes (n =416) on FSDX are targeted by 12 oligonucleotides, three for each strand of each allele, which is approximately double the typical probe coverage for SNPs.\n\nTest interpretationCYCHP files are evaluated by ABMG certified cytogeneticists. Determination of CNVs is consistent with established cytogenetic standards. A minimum of 25-consecutive impacted probes is the baseline determinant for deletions and 50 probes for duplications independent of variant size. Rare CNVs are determined to be \"pathogenic\" if there is sufficient evidence published (at least two independent publications) to indicate that haploinsufficiency or triplosensitivity of the region or gene(s) involved is causative of clinical features or of sufficient overall size (28). If however, there is insufficient but at least preliminary evidence for a causative role for the region or gene(s) therein they are classified as variants of unknown significance (VOUS) independent of CNV size. Areas of absence of heterozygosity (AOH) are also classed as VOUS if of sufficient size and location to increase the risk for conditions with autosomal recessive inheritance or conditions with parent-of-origin/imprinting effects. Other CNVs are typically not reported after determination that they most likely represent normal common population variants and are contained in databases documenting presumptively benign CNVs, e.g., the Database of Genomic Variants (DGV) (29). These parameters were standard independent of the microarray used for analysis in comparative studies.\n\nPublic health importanceA definitive genetic diagnosis facilitates patient access to appropriate and necessary medical and support services. Defining the underlying genetic cause of DD/ID/ASD and/or multiple congenital anomalies (MCA) in each unique patient is vital to understanding etiology, prognosis, and course. It informs physicians of potential comorbid conditions for which a patient should be evaluated and treated proactively and optimally. Improved understandings of the appropriate therapeutic and behavioral approaches to that patient are also enabled. Genetic testing is best provided in the context of an integrated service (30), so FSDX provides comprehensive, clear, readable, and personalized reports for the healthcare provider and a family-friendly section to facilitate understanding of the often-complex results. The report is complemented by availability of pre- and post-test genetic counseling and technical support to providers. Moreover, FSDX includes personalized insurance pre-authorization and appeals assistance to overcome barriers encountered by both providers and families that, in many circumstances, prevent access to crucial genetic testing services (10-11).\n\nPublished reviews, recommendations and guidelinesThe American College of Medical Genetics (ACMG) (2,3), the American Academy of Child and Adolescent Psychiatry (4,7), the American Academy of Pediatrics (5), and the American Academy of Neurology/Child Neurology Society (6) recommend CMA as the first-tier test in the genetic evaluation of children with unexplained DD, ID, or ASD. Considerable data supporting these guidelines are documented in numerous reviews and publications (31-36). The ACMG has also published guidelines on both array design (25) and the validation of arrays, including validation of a new version of a platform in use by the laboratory from the same manufacturer, and of additional sample types (28).

genomics

Detection of epistatic interactions with Random Forest

In order to elucidate the influence of genetic factors on phenotype variation, non-additive genetic interactions (i.e., epistasis) have to be taken into account. However, there is a lack of methods that can reliably detect such interactions, especially for quantitative traits. Random Forest was previously recognized as a powerful tool to identify the genetic variants that regulate trait variation, mainly due to its ability to take epistasis into account. However, although it can account for interactions, it does not specifically detect them. Therefore, we propose three approaches that extract interactions from a Random Forest by testing for specific signatures that arise from interactions, which we termed paired selection frequency, split asymmetry, and selection asymmetry. Since they complement each other for different epistasis types, an ensemble method that combines the three approaches was also created. We evaluated our approaches on multiple simulated scenarios and two different real datasets from different Saccharomyces cerevisiae crosses. We compared them to the commonly used exhaustive pair-wise linear model approach, as well as several two-stage approaches, where loci are pre-selected prior to interaction testing. The Random Forest-based methods presented here generally outperformed the other methods at identifying meaningful genetic interactions both in simulated and real data. Further examination of the results for the simulated and real datasets established how interactions are extracted from the Random Forest, and explained the performance differences between the methods. Thus, the approaches presented here extend the applicability of Random Forest for the genetic mapping of biological traits.\n\nAuthor summaryThe genetic mechanisms underlying biological traits are often complex, involving the effects of multiple genetic variants. Interactions between these variants, also called epistasis, are also common. The machine learning algorithm Random Forest can be used to study genotype-phenotype relationships, by using genetic variants to predict the phenotype. One of Random Forests strengths is its ability to implicitly model interactions. However, Random Forest does not give any information about which predictors specifically interact, i.e. which variants are in epistasis.\n\nHere, we developed three approaches that identify interactions in a Random Forest. We demonstrated their ability to detect genetic interactions using simulations and real data from Saccharomyces cerevisiae. Our Random Forest-based methods generally outperformed several other commonly used approaches at detecting epistasis.\n\nThis study contributes to the long-standing problem of extracting information about the underlying model from a Random Forest. Since Random Forest has many applications outside of genetic association, this work represents a valuable contribution to not only genotype-phenotype mapping research, but also other scientific applications where interactions between predictors in a Random Forest might be of interest.

bioinformatics

Multivariate genome-wide association study of rapid automatized naming and rapid alternating stimulus in Hispanic and African American youth.

Reading disability is a complex neurodevelopmental disorder that is characterized by difficulties in reading despite educational opportunity and normal intelligence. Performance on rapid automatized naming (RAN) and rapid alternating stimulus (RAS) tests gives a reliable predictor of reading outcome. These tasks involve the integration of different neural and cognitive processes required in a mature reading brain. Most studies examining the genetic factors that contribute to RAN and RAS performance have focused on pedigree-based analyses in samples of European descent, with limited representation of groups with Hispanic or African ancestry. In the present study, we conducted a multivariate genome-wide association analysis to identify shared genetic factors that contribute to performance across RAN Objects, RAN Letters, and RAS Letters/Numbers in a sample of Hispanic and African American youth (n=1,331). We then tested whether these factors also contribute to variance in reading fluency and word reading. Genome-wide significant, pleiotropic, effects across RAN Objects, RAN Letters, and RAS Letters/Numbers were observed for SNPs located on chromosome 10q23.31 (rs1555839, multivariate association, p=2.23 x 10-8), which also showed significant association with reading fluency and word reading performance (p <0.001). Bioinformatic analysis of this region using epigenetic data from the NIH Roadmap Epigenomics Mapping Consortium indicates active transcription of the gene RNLS in the brain. Neuroimaging genetic analysis of fourteen cortical regions in an independent sample of typically developing children across multiple ethnicities (n=690) showed that rs1555839 was associated with variation in volume of the right inferior parietal cortex--a region of the brain that processes numerical information and has been implicated in reading disability. This study provides support for a novel locus on chromosome 10q23.31 associated with RAN, RAS, and reading-related performance.\n\nAUTHOR SUMMARYReading disability has a strong genetic component that is explained by multiple genes and genetic factors. The complex genetic architecture along with diverse cognitive impairments associated with reading disability, poses challenges in identifying novel genes and variants that confer risk. One method to begin parsing genetic and neurobiological mechanisms that contribute to reading disability is to take advantage of the high correlation among reading-related cognitive traits like rapid automatized naming (RAN) and rapid alternating stimulus (RAS) to identify shared genetic factors that contribute to common biological mechanisms. In the present study, we used a multivariate genome-wide analysis approach that identified a region of chromosome 10q23.31 associated with variation in RAN Objects, RAN Letters, and RAS Letters/Numbers performance in a sample of 1,331 Hispanic and African American youth in the Genes, Reading, and Dyslexia (GRaD) Study. Genetic variants in this region were also associated with reading fluency in GRaD, and differences in brain structures implicated in reading disability in a separate sample of 690 children. The gene, RNLS, is located within the implicated region of chromosome 10q23.31 and plays a role in breaking down a class of chemical messengers known to affect attention, learning, and memory in the brain. These findings provide a basis to inform our understanding of the biological basis of reading disability.

genetics

Mega-analysis of 31,396 individuals from 6 countries uncovers strong gene-environment interaction for human fertility

Family and twin studies suggest that up to 50% of individual differences in human fertility within a population might be heritable. However, it remains unclear whether the genes associated with fertility outcomes such as number of children ever born (NEB) or age at first birth (AFB) are the same across geographical and historical environments. By not taking this into account, previous genetic studies implicitly assumed that the genetic effects are constant across time and space. We conduct a mega-analysis applying whole genome methods on 31,396 unrelated men and women from six Western countries. Across all individuals and environments, common single-nucleotide polymorphisms (SNPs) explained only ~4% of the variance in NEB and AFB. We then extend these models to test whether genetic effects are shared across different environments or unique to them. For individuals belonging to the same population and demographic cohort (born before or after the 20th century fertility decline), SNP-based heritability was almost five times higher at 22% for NEB and 19% for AFB. We also found no evidence suggesting that genetic effects on fertility are shared across time and space. Our findings imply that the environment strongly modifies genetic effects on the tempo and quantum of fertility, that currently ongoing natural selection is heterogeneous across environments, and that gene-environment interactions may partly account for missing heritability in fertility. Future research needs to combine efforts from genetic research and from the social sciences to better understand human fertility.\n\nAuthors SummaryFertility behavior - such as age at first birth and number of children - varies strongly across historical time and geographical space. Yet, family and twin studies, which suggest that up to 50% of individual differences in fertility are heritable, implicitly assume that the genes important for fertility are the same across both time and space. Using molecular genetic data (SNPs) from over 30,000 unrelated individuals from six different countries, we show that different genes influence fertility in different time periods and different countries, and that the genetic effects consistently related to fertility are presumably small. The fact that genetic effects on fertility appear not to be universal could have tremendous implications for research in the area of reproductive medicine, social science and evolutionary biology alike.

Genetics

Identifying Pleiotropic Effects: A Two-Stage Approach Using Genome-Wide Association Meta-Analysis Data

Pleiotropic effects occur when a single genetic variant independently influences multiple phenotypes. In genetic epidemiological studies, multiple endo-phenotypes or correlated traits are commonly tested separately in a univariate statistical framework to identify associations with genetic determinants. Subsequently, a simple look-up of overlapping univariate results is applied to identify pleiotropic genetic effects. However, this strategy offers limited power to detect pleiotropy. In contrast, combining correlated traits into a composite test provides a powerful approach for detecting pleiotropic genes. Here, we propose a two-stage approach to identify potential pleiotropic effects by utilizing aggregated results from large-scale genome-wide association (GWAS) meta-analyses. In the first stage, we developed two novel approaches (direct linear combining, dLC; and empirical combining, eLC) combining correlated univariate test statistics to screen potential pleiotropic variants on a genome-wide scale, using either individual-level or aggregated data. Our simulations indicated that dLC and eLC outperform other popular multivariate approaches (such as principal component analysis (PCA), multivariate analysis of variance (MANOVA), canonical correlation (CCA), generalized estimation equations (GEE), linear mixed effects models (LME) and OBrien combining approach). In particular, eLC provides a notable increase in power when the genetic variant exhibits both protective and deleterious effects. In the second stage, we developed a unique approach, conditional pleiotropy testing (cPLT), to examine pleiotropic effects using individual-level data for candidate variants identified in Stage 1. Simulation demonstrated reduced type 1 error for cPLT in identifying pleiotropic genetic variants compared to the typical conditional strategy. We validated our two-stage approach by performing a bivariate GWA study on two correlated quantitative traits, high-density lipoprotein (HDL) and triglycerides (TG), in the Genetic Analysis Workshop 16 (GAW16) simulation dataset. In summary, the proposed two-stage approach allows us to leverage aggregated summary statistics from univariate GWAS and improves the power to identify potential pleiotropy while maintaining valid false-positive rates.\n\nAuthor SummaryPleiotropy, occurring when a single genetic variant contributes to multiple phenotypes, remains difficult to identify in genome-wide association studies (GWAS). To leverage data for multiple phenotypes and incorporate univariate GWAS summary results, we propose a novel two-stage approach for discovering potential pleiotropic variants. In the first stage, two novel combining approaches were developed to screen potential pleiotropic variants on a genome-wide scale. Simulations demonstrated the superior statistical power of these approaches over other multivariate methods. In the second stage, our approach was used to identify potential pleiotropy in the candidate marker sets generated from the first stage. The proposed two-stage approach was applied to the GAW16 simulation dataset to discover pleiotropic variants associated with high-density lipoprotein and triglycerides. In summary, we demonstrate that the proposed two-stage approach can be applied as a viable and robust strategy to accommodate phenotypic and genetic heterogeneity for discovering potential pleiotropy on genome-wide scale.

genetics

Adaptive introgression from distant Caribbean islands contributed to the diversification of a microendemic radiation of trophic specialist pupfishes

Rapid diversification often involves complex histories of gene flow that leave variable and conflicting signatures of evolutionary relatedness across the genome. Identifying the extent and source of variation in these evolutionary relationships can provide insight into the evolutionary mechanisms involved in rapid radiations. Here we compare the discordant evolutionary relationships associated with species phenotypes across 42 whole genomes from a sympatric adaptive radiation of Cyprinodon pupfishes endemic to San Salvador Island, Bahamas and several outgroup pupfish species in order to understand the rarity of these trophic specialists within the larger radiation of Cyprinodon. 82% of the genome depicts close evolutionary relationships among the San Salvador Island species reflecting their geographic proximity, but the vast majority of the fixed variants between the specialist species lie in regions with discordant topologies. These regions include signatures of selective sweeps and adaptive introgression from neighboring islands into each of the specialist species. Hard selective sweeps of genetic variation on San Salvador contributed 10-fold more to divergence between specialist species within the radiation than adaptive introgression of Caribbean genetic variation; however, some of these introgressed regions from distant islands were associated with the primary axis of oral jaw divergence within the radiation. For example, standing variation in a proto-oncogene (ski) known to have effects on jaw size introgressed into one San Salvador specialist from an island 300 km away. The complex emerging picture of the origins of adaptive radiation on San Salvador indicates that multiple sources of genetic variation contributed to the adaptive phenotypes of novel trophic specialists on the island. Our findings suggest that a suite of factors, including rare adaptive introgression, may also be required to trigger adaptive radiation in the presence of ecological opportunity.\n\nAuthor summaryGroups of closely related species can rapidly evolve to occupy diverse ecological roles, but the ecological and genetic conditions that trigger this diversification are still highly debated. We examine patterns of molecular evolution across the genomes of a rapid radiation of pupfishes that includes two trophic specialists. Despite apparently widespread ecological opportunities and gene flow across the Caribbean, this radiation is endemic to a single Bahamian Island. Using the whole genomes of 42 pupfish we find evidence of extensive and previously unexpected variation in evolutionary relatedness among Caribbean pupfish. Two sources of genetic variation have contributed to the adaptive diversification of complex phenotypes in this system: selective sweeps of genetic variation from across the Caribbean that was brought into San Salvador through hybridization and genetic variation found on San Salvador. While genetic variation from San Salvador appears to be relatively more common in the divergence observed among specialists, hybridization probably played an important role in the evolution of the complex phenotypes as well. Our findings that multiple sources of genetic variation contribute to the San Salvador radiation suggest that a complex suite of factors, including hybridization with other species, may be required to trigger adaptive radiation in the presence of ecological opportunity.

evolutionary biology

Genome-Wide SNPs Reveal the Drivers of Gene Flow In An Urban Population of the Asian Tiger Mosquito, Aedes albopictus

Aedes albopictus is a highly invasive disease vector with an expanding worldwide distribution. Genetic assays using low to medium resolution markers have found little evidence of spatial genetic structure even at broad geographic scales, suggesting frequent passive movement along human transportation networks. Here we analysed genetic structure of Ae. albopictus collected from 12 sample sites in Guangzhou, China, using thousands of genome-wide single nucleotide polymorphisms (SNPs). We found evidence for passive gene flow, with distance from shipping terminals being the strongest predictor of genetic distance among mosquitoes. As further evidence of passive dispersal, we found multiple pairs of full-siblings distributed between two sample sites 3.7 km apart. After accounting for geographical variability, we also found evidence for isolation by distance, previously undetectable in Ae. albopictus. These findings demonstrate how large SNP datasets and spatially-explicit hypothesis testing can be used to decipher processes at finer geographic scales than formerly possible. Our approach can be used to help predict new invasion pathways of Ae. albopictus and to refine strategies for vector control that involve the transformation or suppression of mosquito populations.\n\nAuthor SummaryAedes albopictus, the Asian Tiger Mosquito, is a highly invasive disease vector with a growing global distribution. Designing strategies to prevent invasion and to control Ae. albopictus populations in invaded regions requires knowledge of how Ae. albopictus disperses. Studies comparing Ae. albopictus populations have found little evidence of genetic structure even between distant populations, suggesting that dispersal along human transportation networks is common. However, a more specific understanding of dispersal processes has been unavailable due to an absence of studies using high-resolution genetic markers. Here we present a study using high-resolution markers, which investigates genetic structure among 152 Ae. albopictus from Guangzhou, China. We found that human transportation networks, particularly shipping terminals, had an influence on genetic structure. We also found genetic distance was correlated with geographical distance, the first such observation in this species. This study demonstrates how high-resolution markers can be used to investigate ecological processes that may otherwise escape detection. We conclude that strategies for controlling Ae. albopictus will have to consider both passive reinvasion along human transportation networks and active reinvasion from neighbouring regions.

ecology

Significant shared heritability underlies suicide attempt and clinically predicted probability of attempting suicide

Suicide accounts for nearly 800,000 deaths per year worldwide with rates of both deaths and attempts rising. Family studies have estimated substantial heritability of suicidal behavior; however, collecting the sample sizes necessary for successful genetic studies has remained a challenge. We utilized two different approaches in independent datasets to characterize the contribution of common genetic variation to suicide attempt. The first is a patient reported suicide attempt phenotype from genotyped samples in the UK Biobank (337,199 participants, 2,433 cases). The second leveraged electronic health record (EHR) data from the Vanderbilt University Medical Center (VUMC, 2.8 million patients, 3,250 cases) and machine learning to derive probabilities of attempting suicide in 24,546 genotyped patients. We identified significant and comparable heritability estimates of suicide attempt from both the patient reported phenotype in the UK Biobank (h2SNP = 0.035, p = 7.12x10-4) and the clinically predicted phenotype from VUMC (h2SNP = 0.046, p = 1.51x10-2). A significant genetic overlap was demonstrated between the two measures of suicide attempt in these independent samples through polygenic risk score analysis (t = 4.02, p = 5.75x10-5) and genetic correlation (rg = 1.073, SE = 0.36, p = 0.003). Finally, we show significant but incomplete genetic correlation of suicide attempt with insomnia (rg = 0.34 - 0.81) as well as several psychiatric disorders (rg = 0.26 - 0.79). This work demonstrates the contribution of common genetic variation to suicide attempt. It points to a genetic underpinning to clinically predicted risk of attempting suicide that is similar to the genetic profile from a patient reported outcome. Lastly, it presents an approach for using EHR data and clinical prediction to generate quantitative measures from binary phenotypes that improved power for our genetic study.

genomics

Polygenic Adaptation: From sweeps to subtle frequency shifts

1Evolutionary theory has produced two conflicting paradigms for the adaptation of a polygenic trait. While population genetics views adaptation as a sequence of selective sweeps at single loci underlying the trait, quantitative genetics posits a collective response, where phenotypic adaptation results from subtle allele frequency shifts at many loci. Yet, a synthesis of these views is largely missing and the population genetic factors that favor each scenario are not well understood. Here, we study the architecture of adaptation of a binary polygenic trait (such as resistance) with negative epistasis among the loci of its basis. The genetic structure of this trait allows for a full range of potential architectures of adaptation, ranging from sweeps to small frequency shifts. By combining computer simulations and a newly devised analytical framework based on Yule branching processes, we gain a detailed understanding of the adaptation dynamics for this trait. Our key analytical result is an expression for the joint distribution of mutant alleles at the end of the adaptive phase. This distribution characterizes the polygenic pattern of adaptation at the underlying genotype when phenotypic adaptation has been accomplished. We find that a single compound parameter, the population-scaled background mutation rate {Theta}bg, explains the main differences among these patterns. For a focal locus, {Theta}bg measures the mutation rate at all redundant loci in its genetic background that offer alternative ways for adaptation. For adaptation starting from mutation-selection-drift balance, we observe different patterns in three parameter regions. Adaptation proceeds by sweeps for small {Theta}bg [prsim] 0.1, while small polygenic allele frequency shifts require large {Theta}bg [scsim] 100. In the large intermediate regime, we observe a heterogeneous pattern of partial sweeps at several interacting loci. 2 Author summaryIt is still an open question how complex traits adapt to new selection pressures. While population genetics champions the search for selective sweeps, quantitative genetics proclaims adaptation via small concerted frequency shifts. To date the empirical evidence of clear sweep signals is more scarce than expected, while subtle shifts remain notoriously hard to detect. In the current study we develop a theoretical framework to predict the expected adaptive architecture of a simple polygenic trait, depending on parameters such as mutation rate, effective population size, size of the trait basis, and the available genetic variability at the onset of selection. For a population in mutation-selection-drift balance we find that adaptation proceeds via complete or partial sweeps for a large set of parameter values. We predict adaptation by small frequency shifts for two main cases. First, for traits with a large mutational target size and high levels of genetic redundancy among loci, and second if the starting frequencies of mutant alleles are more homogeneous than expected in mutation-selection-drift equilibrium, e.g. due to population structure or balancing selection.

evolutionary biology