Search bioRxivSearch

Biology subjects

Evans, D. M.

Publications and source records attributed to Evans, D. M..

12 recordsLinked to original sources

Machine Learning to Predict Osteoporotic Fracture Risk from Genotypes

BackgroundGenomics-based prediction could be useful since genome-wide genotyping costs less than many clinical tests. We tested whether machine learning methods could provide a clinically-relevant genomic prediction of quantitative ultrasound speed of sound (SOS)--a risk factor for osteoporotic fracture.\n\nMethodsWe used 341,449 individuals from UK Biobank with SOS measures to develop genomically-predicted SOS (gSOS) using machine learning algorithms. We selected the optimal algorithm in 5,335 independent individuals and then validated it and its ability to predict incident fracture in an independent test dataset (N = 80,027). Finally, we explored whether genomic pre-screening could complement a UK-based osteoporosis screening strategy, based on the validated tool FRAX.\n\nResultsgSOS explained 4.8-fold more variance in SOS than FRAX clinical risk factors (CRF) alone (r2 = 23% vs. 4.8%). A standard deviation decrease in gSOS, adjusting for the CRF-FRAX score was associated with a higher increased odds of incident major osteoporotic fracture (1,491 cases / 78,536 controls, OR = 1.91 [1.70-2.14], P = 10-28) than that for measured SOS (OR = 1.60 [1.50-1.69], P = 10-52) and femoral neck bone mineral density (147 cases / 4,594 controls, OR = 1.53 [1.27-1.83], P = 10-6). Individuals in the bottom decile of the gSOS distribution had a 3.25-fold increased risk of major osteoporotic fracture (P = 10-18) compared to the top decile. A gSOS-based FRAX score, identified individuals at high risk for incident major osteoporotic fractures better than the CRF-FRAX score (P = 10-14). Introducing a genomic pre-screening step into osteoporosis screening in 4,741 individuals reduced the number of required clinical visits from 2,455 to 1,273 and the number of BMD tests from 1,013 to 473, while only reducing the sensitivity to identify individuals eligible for therapy from 99% to 95%.\n\nInterpretationThe use of genotypes in a machine learning algorithm resulted in a clinically-relevant prediction of SOS and fracture, with potential to impact healthcare resource utilization.\n\nResearch in ContextO_ST_ABSEvidence Before this StudyC_ST_ABSGenome-wide association studies have identified many loci associated with risk of clinically-relevant fracture risk factors, such as SOS. Yet, it is unclear if such information can be leveraged to identify those at risk for disease outcomes, such as osteoporotic fractures. Most previous attempts to predict disease risk from genotypes have used polygenic risk scores, which may not be optimal for genomic-prediction. Despite these obstacles, genomic-prediction could enable screening programs to be more efficient since most people screened in a population are not determined to have a level of risk that would prompt a change in clinical care. Genomic pre-screening could help identify individuals whose risk of disease is low enough that they are unlikely to benefit from screening.\n\nAdded Value of this StudyUsing a large dataset of 426,811 individuals we trained and tested a machine learning algorithm to genomically-predict SOS. This metric, gSOS, had performance characteristics for predicting fracture risk that were similar to measured SOS and femoral neck BMD. Implementing a gSOS-based pre-screening step into the UK-based osteoporosis treatment guidelines reduced the number of individuals who would require screening clinical visits and skeletal testing by approximately 50%, while having little impact on the sensitivity to identify individuals at high risk for osteoporotic fracture.\n\nImplications of all of the Available EvidenceClinically-relevant genomic prediction of heritable traits is feasible using the machine learning algorithm presented here in large sample sizes. Genome-wide genotyping is now less expensive than many clinical tests, needs to be performed once over a lifetime and could risk stratify for multiple heritable traits and diseases years prior to disease onset, providing an opportunity for prevention. The implementation of such algorithms could improve screening efficiency, yet their cost-effectiveness will need to be ascertained in subsequent analyses.

genomics

Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: a comparison using microscopy and DNA metabarcoding

O_LIMoths are globally relevant as pollinators but nocturnal pollination remains poorly understood. Plant-pollinator interaction networks are traditionally constructed using either flower-visitor observations or pollen-transport detection using microscopy. Recent studies have shown the potential of DNA metabarcoding for detecting and identifying pollen-transport interactions. However, no study has directly compared the realised observations of pollen-transport networks between DNA metabarcoding and conventional light microscopy.\nC_LIO_LIUsing matched samples of nocturnal moths, we construct pollen-transport networks using two methods: light microscopy and DNA metabarcoding. Focussing on the feeding mouthparts of moths, we develop and provide reproducible methods for merging DNA metabarcoding and ecological network analysis to better understand species-interactions.\nC_LIO_LIDNA metabarcoding detected pollen on more individual moths, and detected multiple pollen types on more individuals than microscopy, but the average number of pollen types per individual was unchanged. However, after aggregating individuals of each species, metabarcoding detected more interactions per moth species. Pollen-transport network metrics differed between methods, because of variation in the ability of each to detect multiple pollen types per moth and to separate morphologically-similar or related pollen. We detected unexpected but plausible moth-plant interactions with metabarcoding, revealing new detail about nocturnal pollination systems.\nC_LIO_LIThe nocturnal pollination networks observed using metabarcoding and microscopy were similar, yet distinct, with implications for network ecologists. Comparisons between networks constructed using metabarcoding and traditional methods should therefore be treated with caution. Nevertheless, the potential applications of metabarcoding for studying plant-pollinator interaction networks are encouraging, especially when investigating understudied pollinators such as moths.\nC_LI

ecology

Circulating selenium and prostate cancer risk: a Mendelian randomization analysis

In the Selenium and Vitamin E Cancer Prevention Trial (SELECT), selenium supplementation (causing a median 114 g/L increase in circulating selenium) did not lower overall prostate cancer risk, but increased risk of high-grade prostate cancer and type 2 diabetes. Mendelian randomization analysis uses genetic variants to proxy modifiable risk factors and can strengthen causal inference in observational studies. We constructed a genetic risk score comprising eleven single-nucleotide polymorphisms robustly (P<5x10-8) associated with circulating selenium in genome-wide association studies. In a Mendelian randomization analysis of 72,729 men in the PRACTICAL Consortium (44,825 cases, 27,904 controls), 114 g/L higher genetically-elevated circulating selenium was not associated with prostate cancer (OR: 1.01; 95% CI: 0.89-1.13). Concordant with findings from SELECT, selenium was weakly associated with advanced (including high-grade) prostate cancer (OR: 1.21; 95% CI: 0.98-1.49) and type 2 diabetes (OR: 1.18; 95% CI: 0.97-1.43; in a type 2 diabetes GWAS meta-analysis with up to 49,266 cases, 249,906 controls). Mendelian randomization mirrored the outcome of selenium supplementation in SELECT and may offer an approach for the prioritization of interventions for follow-up in large-scale randomized controlled trials.

epidemiology

Estimating sampling completeness of interactions in quantitative bipartite ecological networks: incorporating variation in species specialisation

BackgroundThe analysis of ecological networks can be affected by sampling effort, potentially leading to bias. Ecological network structure is often summarised by descriptive metrics but these metrics can vary according to the proportion of the total interactions that have been observed. Therefore, to know the likely degree of bias, it is valuable to estimate the total number of interactions in a network, and so calculate the proportion of interactions that have been observed (sampling completeness of interactions). Existing approaches to estimate sampling completeness of interactions use the Chao family of asymptotic species richness estimators to predict the total number of interactions, but do not fully utilise information about the relative specialisation of species within the network.\n\nResultsHere, we propose a modification of previously-used methods, that places equal weight on each interaction (whether or not it has been observed), rather than on each species. Our approach is therefore equivalent to weighting the interaction sampling completeness of each species in the network according to its relative specialisation. We demonstrate that, for the subset of species that are observed and when assuming that species richness estimators accurately project the number of unobserved interactions per observed species, our approach is mathematically more accurate. Our approach can be universally applied to any quantitative, bipartite network.\n\nWe propose two methods to estimation using our approach, using abundance-based and incidence-based species richness estimators respectively, and give recommendations when each should be applied. We discuss the effect of unobserved species and the potential use of a threshold of minimum abundance for species inclusion. Finally, we consider these advances in the context of some of the main issues surrounding estimation of interaction sampling completeness in network ecology.\n\nConclusionsWe recommend that future studies of bipartite networks utilise our approach and methods to estimate the sampling completeness of interactions, to assist with the quantitative and comparative analysis and interpretation of network properties.

ecology

Developmental changes within the genetic architecture of social communication behaviour: A multivariate study of genetic variance in unrelated individuals

BackgroundRecent analyses of trait-disorder overlap suggest that psychiatric dimensions may relate to distinct sets of genes that exert their maximum influence during different periods of development. This includes analyses of social-communciation difficulties that share, depending on their developmental stage, stronger genetic links with either Autism Spectrum Disorder or schizophrenia. Here we developed a multivariate analysis framework in unrelated individuals to model directly the developmental profile of genetic influences contributing to complex traits, such as social-communication difficulties, during a [~]10-year period spanning childhood and adolescence.\n\nMethodsLongitudinally assessed quantitative social-communication problems (N[&le;] 5,551) were studied in participants from a UK birth cohort (ALSPAC, 8 to 17 years). Using standardised measures, genetic architectures were investigated with novel multivariate genetic-relationship-matrix structural equation models (GSEM) incorporating whole-genome genotyping information. Analogous to twin research, GSEM included Cholesky decomposition, common pathway and independent pathway models.\n\nResultsA 2-factor Cholesky decomposition model described the data best. One genetic factor was common to SCDC measures across development, the other accounted for independent variation at 11 years and later, consistent with distinct developmental profiles in trait-disorder overlap. Importantly, genetic factors operating at 8 years explained only [~]50% of the genetic variation at 17 years.\n\nConclusionUsing latent factor models, we identified developmental changes in the genetic architecture of social-communication difficulties that enhance the understanding of ASD and schizophrenia-related dimensions. More generally, GSEM present a framework for modelling shared genetic aetiologies between phenotypes and can provide prior information with respect to patterns and continuity of trait-disorder overlap.

genetics

Novel pleiotropic risk loci for melanoma and nevus density implicate multiple biological pathways

The total number of acquired melanocytic nevi on the skin is strongly correlated with melanoma risk. Here we report a meta-analysis of 11 nevus GWAS from Australia, Netherlands, United Kingdom, and United States, comprising a total of 52,506 phenotyped individuals. We confirm known loci including MTAP, PLA2G6, and IRF4, and detect novel SNPs at a genome-wide level of significance in KITLG, DOCK8, and a broad region of 9q32. In a bivariate analysis combining the nevus results with those from a recent melanoma GWAS meta-analysis (12,874 cases, 23,203 controls), SNPs near GPRC5A, CYP1B1, PPARGC1B, HDAC4, FAM208B and SYNE2 reached global significance, and other loci, including MIR146A and OBFC1, reached a suggestive level of significance. Overall, we conclude that most nevus genes affect melanoma risk (KITLG an exception), while many melanoma risk loci do not alter nevus count. For example, variants in TERC and OBFC1 affect both traits, but other telomere length maintenance genes seem to affect melanoma risk only. Our findings implicate multiple pathways in nevogenesis via genes we can show to be expressed under control of the MITF melanocytic cell lineage regulator.

genetics

Using Structural Equation Modeling to Jointly Estimate Maternal and Foetal Effects on Birthweight in the UK Biobank

BackgroundTo date, 60 genetic variants have been robustly associated with birthweight. It is unclear whether these associations represent the effect of an individuals own genotype on their birthweight, their mothers genotype, or both.\n\nMethodsWe demonstrate how structural equation modelling (SEM) can be used to estimate both maternal and foetal effects when phenotype information is present for individuals in two generations and genotype information is available on the older individual. We conduct an extensive simulation study to assess the bias, power and type 1 error rates of the SEM and also apply the SEM to birthweight data in the UK Biobank study.\n\nResultsUnlike simple regression models, our approach is unbiased when there is both a maternal and foetal effect. The method can be used when either the individuals own phenotype or the phenotype of their offspring is not available, and allows the inclusion of summary statistics from additional cohorts where raw data cannot be shared. We show that the type 1 error rate of the method is appropriate, there is substantial statistical power to detect a genetic variant that has a moderate effect on the phenotype, and reasonable power to detect whether it is a foetal and/or maternal effect. We also identify a subset of birth weight associated SNPs that have opposing maternal and foetal effects in the UK Biobank.\n\nConclusionsOur results show that SEM can be used to estimate parameters that would be difficult to quantify using simple statistical methods alone.\n\nKey MessagesO_LIWe describe a structural equation model to estimate both maternal and foetal effects when phenotype information is present for individuals in two generations and genotype information is available on the older individual.\nC_LIO_LIUsing simulation, we show that our approach is unbiased when there is both a maternal and foetal effect, unlike simple linear regression models. Additionally, we illustrate that the structural equation model is largely robust to measurement error and missing data for either the individuals own phenotype or the phenotype of their offspring.\nC_LIO_LIWe describe how the flexibility of the structural equation modelling framework will allow the inclusion of summary statistics from studies that are unable to share raw data.\nC_LIO_LIUsing the structural equation model to estimate the maternal and foetal effects of known birthweight associated loci in the UK Biobank, we identify three loci that have primary effects through the maternal genome and six loci that have opposite effects in the maternal and foetal genomes.\nC_LI

genetics

Partitioning Phenotypic Variance Due To Parent-Of-Origin Effects Using Genomic Relatedness Matrices

Introduction Introduction Methods Statistical Methods Results Discussion References Parent-of-origin effects (POEs) describe the phenomenon in which the effects of alleles depend upon their parental origin. POEs imply that heterozygote individuals have phenotypes which are distributed differently depending upon which of their alleles were maternally and paternally transmitted (Guilmatre and Sharp 2012; Lawson et al. 2013). The extreme case of POEs is polar overdominance, where the two heterozygotes' phenotypes differ in distribution but the two homozygotes share the same distribution (Hoggart et al. 2014). Imprinting, a phenomenon in which one parent's allele is not expressed, is probably the most widely studied example of POE (Peters 2 ...

genetics

Diagnostic Yield And Treatment Impact Of Targeted Exome Sequencing In Early-Onset Epilepsy

BackgroundTo examine the impact on diagnosis, treatment and cost with early use of targeted whole-exome sequencing (WES) in early-onset epilepsy.\n\nMethodsWES was performed on 50 patients with early-onset epilepsy ([&le;] 5 years) of unknown cause. Patients were classified as retrospective (epilepsy diagnosis > 6 months) or prospective (epilepsy diagnosis < 6 months). WES was performed on an Ion ProtonTM and variant reporting was restricted to the sequences of 565 known epilepsy genes. Diagnostic yield and time to diagnosis were calculated. An analysis of cost and impact on treatment was also performed.\n\nResultsA likely/definite diagnosis was made in 17/50 patients (34%) with immediate treatment implications in 8/17 (47%). A possible diagnosis was identified in 9 additional patients (18%) for whom supporting evidence is pending. Time from epilepsy onset to genetic diagnosis was faster when WES was performed early in the diagnostic process (mean: 143 days prospective versus 2,172 days retrospective). Costs of prior negative tests averaged $8,344 in the retrospective group, suggesting savings of up to $5,110 per patient.\n\nInterpretationThese results support the clinical utility and potential cost-effectiveness of using targeted WES early in the diagnostic workup of patients with unexplained early-onset epilepsy. The costs and clinical benefits are likely to continue to improve. Advances in precision medicine and further studies regarding impact on long-term clinical outcome will be important.

genetics

Causal Analyses, Statistical Efficiency And Phenotypic Precision Through Recall-By-Genotype Study Design

Genome-wide association studies have been useful in identifying common genetic variants related to a variety of complex traits and diseases; however, they are often limited in their ability to inform about underlying biology. Whilst bioinformatics analyses, studies of cells, animal models and applied genetic epidemiology have provided some understanding of genetic associations or causal pathways, there is a need for new genetic studies that elucidate causal relationships and mechanisms in a cost-effective, precise and statistically efficient fashion. We discuss the motivation for and the characteristics of the Recall-by-Genotype (RbG) study design, an approach that enables genotype-directed deep-phenotyping and improvement in drawing causal inferences. Specifically, we present RbG designs using single and multiple variants and discuss the inferential properties, analytical approaches and applications of both. We consider the efficiency of the RbG approach, the likely value of RbG studies for the causal investigation of disease aetiology and the practicalities of incorporating genotypic data into population studies in the context of the RbG study design. Finally, we provide a catalogue of the UK-based resources for such studies, an online tool to aid the design of new RbG studies and discuss future developments of this approach.

genetics

Testing the principles of Mendelian randomization: Opportunities and complications on a genomewide scale

BackgroundMendelian randomization (MR) uses genetic variants as instrumental variables to assess whether observational associations between exposures and disease reflect causal relationships. MR requires genetic variants to be independent of factors that confound observational associations.\n\nMethodsUsing data from the Avon Longitudinal Study of Parents and Children, associations within and between 121 phenotypes and 13,720 genetic variants (from the NHGRI-EBI GWAS catalog) were examined to assess the validity of MR assumptions.\n\nResultsAmongst 7,260 pairwise comparisons between the 121 phenotypes, 2,188 (30%) provided evidence of association, where 363 were expected at the 5% level (observed:expected ratio=6.03; 95% CI: 5.42, 6.70; {chi}2=9682.29; d.f. =1, P[&le;]1x10-50). Amongst 1,660,120 pairwise associations between phenotypes and genotypes, 86,748 (5.2%) gave evidence of association at the same threshold, where 83,006 were expected (observed:expected ratio=1.05; 95% CI: 1.04, 1.05; {chi}2=117.57; d.f. =1, P=2.15x10-27). Amongst 1,171,764 pairwise associations between the phenotypes and LD pruned independent genetic variants, 60,136 (5.1%) gave evidence of association, where 58,588 were expected (observed:expected ratio=1.03; 95% CI: 1.03, 1.08; {chi}2= 43.05; d.f. = 1, P=5.33x10-11).\n\nConclusionThese results confirm previously observed patterns of phenotypic correlation. They also provide evidence of a substantially lower level of association between genetic variants and phenotypes, with residual inflation the likely product of indistinguishable real genetic association, multiple variables measuring the same biological phenomena, or pleiotropy. These results reflect the favorable properties of genetic instruments for estimating causal relationships, but confirm the need for functional information or analytical methods to account for pleiotropic events.

epidemiology

MR-Base: a platform for systematic causal inference across the phenome using billions of genetic associations

Published genetic associations can be used to infer causal relationships between phenotypes, bypassing the need for individual-level genotype or phenotype data. We have curated complete summary data from 1094 genome-wide association studies (GWAS) on diseases and other complex traits into a centralised database, and developed an analytical platform that uses these data to perform Mendelian randomization (MR) tests and sensitivity analyses (MR-Base, http://www.mrbase.org). Combined with curated data of published GWAS hits for phenomic measures, the MR-Base platform enables millions of potential causal relationships to be evaluated. We use the platform to predict the impact of lipid lowering on human health. While our analysis provides evidence that reducing LDL-cholesterol, lipoprotein(a) or triglyceride levels reduce coronary disease risk, it also suggests causal effects on a number of other non-vascular outcomes, indicating potential for adverse-effects or drug repositioning of lipid-lowering therapies.

epidemiology