Search bioRxivSearch

Biology subjects

Michael Inouye

Publications and source records attributed to Michael Inouye.

15 recordsLinked to original sources

Identification of Brain Expression Quantitative Trait Loci Associated with Schizophrenia and Affective Disorders in Normal Brain Tissue

Schizophrenia and the affective disorders, here comprising bipolar disorder and major depressive disorder, are psychiatric illnesses that lead to significant morbidity and mortality worldwide. Whilst understanding of their pathobiology remains limited, large case-control studies have recently identified single nucleotide polymorphisms (SNPs) associated with these disorders. However, discerning the functional effects of these SNPs has been difficult as the associated causal genes are unknown. Here we evaluated whether schizophrenia and affective disorder associated-SNPs are correlated with gene expression within human brain tissue. Specifically, to identify expression quantitative trait loci (eQTLs), we leveraged disorder-associated SNPs identified from six Psychiatric Genomics Consortium and CONVERGE Consortium studies with gene expression levels in post-mortem, neurologically-normal tissue from two independent human brain tissue expression datasets (UK Brain Expression Consortium (UKBEC) and Genotype-Tissue Expression (GTEx)). We identified 6 188 and 16 720 cis-acting SNPs exceeding genome-wide significance (p<5x10-8) in the UKBEC and GTEx datasets, respectively. 1 288 cis-eQTLs were significant in a metaanalysis leveraging overlapping brain regions and were associated with expression of 15 genes, including three non-coding RNAs. One cis-eQTL, rs 16969968, results in a functionally disruptive missense mutation in CHRNA5, a schizophrenia-implicated gene. Meta-analysis identified 297 trans-eQTLs associated with 24 genes that were significant in a region-specific manner. Importantly, comparing across tissues, we find that blood eQTLs largely do not capture brain cis-eQTLs. This study identifies putatively causal genes whose expression in region-specific brain tissue may contribute to the risk of schizophrenia and affective disorders.

Genetics

Genetic loci associated with coronary artery disease harbor evidence of selection and antagonistic pleiotropy

Traditional genome-wide scans for positive selection have mainly uncovered selective sweeps associated with monogenic traits. While selection on quantitative traits is much more common, very few signals have been detected because of their polygenic nature. We searched for positive selection signals underlying coronary artery disease (CAD) in worldwide populations, using novel approaches to quantify relationships between polygenic selection signals and CAD genetic risk. We identified new candidate adaptive loci that appear to have been directly modified by disease pressures given their significant associations with CAD genetic risk. These candidates were all uniquely and consistently associated with many different male and female reproductive traits suggesting selection may have also targeted these because of their direct effects on fitness. This suggests the presence of widespread antagonistic-pleiotropic tradeoffs on CAD loci, which provides a novel explanation for the maintenance and high prevalence of CAD in modern humans. Lastly, we found that positive selection more often targeted CAD gene regulatory variants using HapMap3 lymphoblastoid cell lines, which further highlights the unique biological significance of candidate adaptive loci underlying CAD. Our study provides a novel approach for detecting selection on polygenic traits and evidence that modern human genomes have evolved in response to CAD-induced selection pressures and other early-life traits sharing pleiotropic links with CAD.\n\nAuthor SummaryHow genetic variation contributes to disease is complex, especially for those such as coronary artery disease (CAD) that develop over the lifetime of individuals. One of the fundamental questions about CAD -- whose progression begins in young adults with arterial plaque accumulation leading to life-threatening outcomes later in life -- is why natural selection has not removed or reduced this costly disease. It is the leading cause of death worldwide and has been present in human populations for thousands of years, implying considerable pressures that natural selection should have operated on. Our study provides new evidence that genes underlying CAD have recently been modified by natural selection and that these same genes uniquely and extensively contribute to human reproduction, which suggests that natural selection may have maintained genetic variation contributing to CAD because of its beneficial effects on fitness. This study provides novel evidence that CAD has been maintained in modern humans as a byproduct of the fitness advantages those genes provide early in human lifecycles.

Genomics

FlashPCA: fast sparse canonical correlation analysis of genomic data

SummarySparse canonical correlation analysis (SCCA) is a useful approach for correlating one set of measurements, such as single nucleotide polymorphisms (SNPs), with another set of measurements, such as gene expression levels. We present a fast implementation of SCCA, enabling rapid analysis of hundreds of thousands of SNPs together with thousands of phenotypes. Our approach is implemented both as an R package flashpcaR and within the standalone commandline tool flashpca.\n\nAvailability and implementationhttps://github.com/gabraham/flashpca\n\nContactgad.abraham@unimelb.edu.au

Genomics

Exploratory analysis and error modeling of a sequencing technology

Next generation DNA sequencing methods have created an unprecedented leap in sequence data generation, thus novel computational tools and statistical models are required to optimize and assess the resulting data. In this report, we explore underlying causes of error for the Illumina Genome Analyzer (IGA) sequencing technology and attempt to quantify their effects using a human bacterial artificial chromosome sequenced to 60,000 fold coverage. Seven potential error predictors are considered: Phred score, read entropy, tile coordinates, local tile density, base position within read, nucleotide call, and lane. With these parameters, logistic regression and log-linear models are constructed and used to show that each of the potential predictors contributes to error (P<1x10-4). With this additional information, we apply the logistic model and achieve a 3% improvement in both the sensitivity and specificity to detect IGA errors. Further, we demonstrate that these modeling approaches can be used as a feedback loop to inform laboratory methods and identify specific machine or run bias.

Genomics

Genomic prediction of coronary heart disease

BackgroundGenetics plays an important role in coronary heart disease (CHD) but the clinical utility of a genomic risk score (GRS) relative to clinical risk scores, such as the Framingham Risk Score (FRS), is unclear.\n\nMethodsWe generated a GRS of 49,310 SNPs based on a CARDIoGRAMplusC4D Consortium meta-analysis of CHD, then independently tested this using five prospective population cohorts (three FINRISK cohorts, combined n=12,676, 757 incident CHD events; two Framingham Heart Study cohorts (FHS), combined n=3,406, 587 incident CHD events).\n\nResultsThe GRS was strongly associated with time to CHD event (FINRISK HR=1.74, 95% CI 1.61-1.86 per S.D. of GRS; Framingham HR=1.28, 95% CI 1.18-1.38), and was largely unchanged by adjustment for clinical risk scores or individual risk factors, including family history. Integration of the GRS with clinical risk scores (FRS and ACC/AHA13 score) improved prediction of CHD events within 10 years (meta-analysis C-index: +1.5-1.6%, P<0.001), particularly for individuals [&ge;]60 years old (meta-analysis C-index: +4.6-5.1%, P<0.001). Men in the top 20% of the GRS had 3-fold higher risk of CHD by age 75 in FINRISK and 2-fold in FHS, and attaining 10% cumulative CHD risk 18y earlier in FINRISK and 12y earlier in FHS than those in the bottom 20%. Furthermore, high genomic risk was partially compensated for by low systolic blood pressure, low cholesterol level, and non-smoking.\n\nConclusionsA GRS based on a large number of SNPs substantially improves CHD risk prediction and encodes decades of variation in CHD risk not captured by traditional clinical risk scores.

Genomics

Mergeomics: integration of diverse genomics resources to identify pathogenic perturbations to biological systems

Mergeomics is a computational pipeline (http://mergeomics.research.idre.ucla.edu/Download/Package/) that integrates multidimensional omics-disease associations, functional genomics, canonical pathways and gene-gene interaction networks to generate mechanistic hypotheses. It first identifies biological pathways and tissue-specific gene subnetworks that are perturbed by disease-associated molecular entities. The disease-associated subnetworks are then projected onto tissue-specific gene-gene interaction networks to identify local hubs as potential key drivers of pathological perturbations. The pipeline is modular and can be applied across species and platform boundaries, and uniquely conducts pathway/network level meta-analysis of multiple genomic studies of various data types. Application of Mergeomics to cholesterol datasets revealed novel regulators of cholesterol metabolism.

Systems Biology

EcOH: In silico serotyping of E. coli from short read data

The lipopolysaccharide (O) and flagellar (H) surface antigens of Escherichia coli are targets for serotyping that have traditionally been used to identify pathogenic lineages of E. coli. As serotyping has several limitations, public health reference laboratories are increasingly moving towards whole genome sequencing (WGS) for the rapid characterisation of bacterial isolates. Here we present a method to rapidly and accurately serotype E. coli isolates from raw, short read sequence data, leveraging the known genetic basis for the biosynthesis of O- and H-antigens. Our approach bypasses the need for de novo genome assembly by directly screening WGS reads against a curated database of alleles linked to known E. coli O-groups and H-types (the EcOH database) using the software package SRST2. We validated our approach by comparing in silico results with those obtained via serological phenotyping of 197 enteropathogenic (EPEC) isolates. We also demonstrated the utility of our method to characterise enterotoxigenic E. coli (ETEC) and the uropathogenic E. coli (UPEC) epidemic clone ST131, and for in silico serotyping of foodborne outbreak-related isolates in the public GenomeTrakr database.

Microbiology

A scalable permutation approach reveals replication and preservation patterns of gene coexpression modules

Gene coexpression network modules provide a framework for identifying shared biological functions. Analysis of topological preservation of modules across datasets is important for assessing reproducibility, and can reveal common function between tissues, cell types, and species. Although module preservation statistics have been developed, heuristics have been required for significance testing. However, the scale of current and future analyses requires accurate and unbiased p-values, particularly to address the challenge of multiple testing. Here, we developed a rapid and efficient approach (NetRep) for assessing module preservation and show that module preservation statistics are typically non-normal, necessitating a permutation approach. Quantification of module preservation across brain, liver, adipose, and muscle tissues in a BxH mouse cross revealed complex patterns of multi-tissue preservation with 52% of modules showing unambiguous preservation in one or more tissues and 25% showing preservation in all four tissues. Phenotype association analysis uncovered a liver-derived gene module which harboured housekeeping genes and which also displayed adipose and muscle tissue specific association with body weight. Taken together, our study presents a rapid unbiased approach for testing preservation of gene network topology, thus enabling rigorous assessment of potentially conserved function and phenotype association analysis.

Bioinformatics

Systems medicine links microbial inflammatory response with glycoprotein-associated mortality risk

Integrative analyses of high-throughput omics data have elucidated the aetiology and pathogenesis for complex traits and diseases1-4, and the linking of omics information to electronic health records promises new insights into human health and disease. Recent nuclear magnetic resonance (NMR) spectroscopy biomarker profiling has implicated glycoprotein acetyls (GlycA) as a biomarker for cardiovascular risk5 and all-cause mortality6. To elucidate biological processes contributing to GlycA-associated mortality risk, we leveraged human omics data from three population-based cohorts together with nation-wide Finnish hospital and mortality records. Elevated GlycA was associated with myriad infection-related inflammatory processes. Within individuals, elevated GlycA levels were stable over long time periods, up to a decade, and chronically elevated GlycA was also associated with modest elevation of numerous cytokines. Individuals with elevated GlycA also showed increased expression of a transcriptional sub-network, the Neutrophil Degranulation Module (NDM), suggesting an increased activity of microbe-driven immune response. Subsequent analysis of nation-wide hospitalisation and death records was consistent with a microbial basis for GlycA-associated mortality, with each standard deviation increase in GlycA raising an individuals future risk of hospitalization and death from non-localized infection by 40% and 136%, respectively. These results show that, beyond its established role in acute-phase response7-9, elevated GlycA is more broadly a biomarker for low-grade chronic inflammation and increased neutrophil activity. Further, increased risk of susceptibility to severe microbial-infection events in healthy individuals suggests this inflammation is a contributor to mortality risk. Taken together, this study demonstrates the power of an integrative approach that combines omics data and health records to delineate the biological processes underlying a newly discovered biomarker, providing a model strategy for future systems medicine studies.

Genomics

Genomic prediction of celiac disease targeting HLA-positive individuals

BackgroundGenomic prediction aims to leverage genome-wide genetic data towards better disease diagnostics and risk scores. We have previously published a genomic risk score (GRS) for celiac disease (CD), a common and highly heritable autoimmune disease, which differentiates between CD cases and population-based controls at a clinically-relevant predictive level, improving upon other gene-based approaches. HLA risk haplotypes, particularly HLA-DQ2.5, are necessary but not sufficient for CD, with at least one HLA risk haplotype present in up to half of most Caucasian populations. Here, we assess a genomic prediction strategy that specifically targets this common genetic susceptibility subtype, utilizing a supervised learning procedure for CD that leverages known HLA-DQ2.5 risk.\n\nMethodsUsing L1/L2-regularized support-vector machines trained on large European case-control datasets, we constructed novel CD GRSs specific to individuals with HLA-DQ2.5 risk haplotypes (GRS-DQ2.5) and compared them with the predictive power of the existing CD GRS (GRS14) as well as two haplotype-based approaches, externally validating the results in a North American case-control study.\n\nResultsConsistent with previous observations, both the existing GRS14 and the GRS-DQ2.5 had better predictive performance than the HLA haplotype approaches. GRS-DQ2.5 models, based on directly genotyped or imputed markers, achieved similar levels of predictive performance (AUC = 0.718--0.73), which were substantially higher than those obtained from the DQ2.5 zygosity alone (AUC = 0.558), the HLA risk haplotype method (AUC = 0.634), or the generic GRS14 (AUC = 0.679). In a screening model of at-risk individuals, the GRS-DQ2.5 lowered the number of unnecessary follow-up tests for CD across most sensitivity levels. Relative to a baseline implicating all DQ2.5-positive individuals for follow-up, the GRS-DQ2.5 resulted in a net saving of 2.2 unnecessary follow-up tests for each justified test while still capturing 90% of DQ2.5-positive CD cases.\n\nConclusionsGenomic risk scores for CD that target genetically at-risk sub-groups improve predictive performance beyond traditional approaches and may represent a useful strategy for prioritizing individuals at increase risk of disease, thus potentially reducing unnecessary follow-up diagnostic tests.

Genomics

The infant airway microbiome in health and disease impacts later asthma development

The nasopharynx (NP) is a reservoir for microbes associated with acute respiratory illnesses (ARI). The development of asthma is initiated during infancy, driven by airway inflammation associated with infections. Here, we report viral and bacterial community profiling of NP aspirates across a birth cohort, capturing all lower respiratory illnesses during their first year. Most infants were initially colonized with Staphylococcus or Corynebacterium before stable colonization with Alloiococcus or Moraxella, with transient incursions of Streptococcus, Moraxella or Haemophilus marking virus-associated ARIs. Our data identify the NP microbiome as a determinant for infection spread to the lower airways, severity of accompanying inflammatory symptoms, and risk for future asthma development. Early asymptomatic colonization with Streptococcus was a strong asthma predictor, and antibiotic usage disrupted asymptomatic colonization patterns.

Genomics

SRST2: Rapid genomic surveillance for public health and hospital microbiology labs

Rapid molecular typing of bacterial pathogens is critical for public health epidemiology, surveillance and infection control, yet routine use of whole genome sequencing (WGS) for these purposes poses significant challenges. Here we present SRST2, a read mapping-based tool for fast and accurate detection of genes, alleles and multi-locus sequence types (MLST) from WGS data. Using >900 genomes from common pathogens, we show SRST2 is highly accurate and outperforms assembly-based methods in terms of both gene detection and allele assignment. Here we have demonstrated the use of SRST2 for microbial genome surveillance in a variety of public health and hospital settings. In the face of rising threats of antimicrobial resistance and emerging virulence amongst bacterial pathogens, SRST2 represents a powerful tool for rapidly extracting clinically useful information from raw WGS data. Source code is available from http://katholt.github.io/srst2/.

Genomics

Towards a molecular systems model of coronary artery disease

Coronary artery disease (CAD) is a complex disease driven by myriad interactions of genetics and environmental factors. Traditionally, studies have analyzed only one disease factor at a time, providing useful but limited understanding of the underlying etiology. Recent advances in cost-effective and high-throughput technologies, such as single nucleotide polymorphism (SNP) genotyping, exome/genome sequencing, gene expression microarrays and metabolomics assays have enabled the collection of millions of data points in many thousands of individuals. In order to make sense of such omics data, effective analytical methods are needed. We review and highlight some of the main results in this area, focusing on integrative approaches that consider multiple modalities simultaneously. Such analyses have the potential to uncover the genetic basis of CAD, produce genomic risk scores (GRS) for disease prediction, disentangle the complex interactions underlying disease, and predict response to treatment.

Systems Biology

Epistasis within the MHC contributes to the genetic architecture of celiac disease

Epistasis has long been thought to contribute to the genetic aetiology of complex diseases, yet few robust epistatic interactions in humans have been detected. We have conducted exhaustive genome-wide scans for pairwise epistasis in five independent celiac disease (CD) case-control studies, using a rapid model-free approach to examine over 500 billion SNP pairs in total. We found 20 significant epistatic signals within the HLA region which achieved stringent replication criteria across multiple studies and were independent of known CD risk HLA haplotypes. The strongest independent CD epistatic signal corresponded to genes in the HLA class III region, in particular PRRC2A and GPANK1/C6orf47, which are known to contain variants for non-Hodgkins lymphoma and early menopause, co-morbidities of celiac disease. Replicable evidence for epistatic variants outside the MHC was not observed. Both within and between European populations, we observed striking consistency of epistatic models and epistatic model distribution. Within the UK population, models of CD based on both epistatic and additive single-SNP effects increased explained CD variance by approximately 1% over those of single SNPs. Models of only epistatic pairs or additive single-SNPs showed similar levels of CD variance explained, indicating the existence of a substantial overlap of additive and epistatic components. Our findings have implications for the determination of genetic architecture and, by extension, the use of human genetics for validation of therapeutic targets.

Genomics

Fast Principal Component Analysis of Large-Scale Genome-Wide Data

Principal component analysis (PCA) is routinely used to analyze genome-wide single-nucleotide polymorphism (SNP) data, for detecting population structure and potential outliers. However, the size of SNP datasets has increased immensely in recent years and PCA of large datasets has become a time consuming task. We have developed flashpca, a highly efficient PCA implementation based on randomized algorithms, which delivers identical accuracy in extracting the top principal components compared with existing tools, in substantially less time. We demonstrate the utility of flashpca on both HapMap3 and on a large Immunochip dataset. For the latter, flashpca performed PCA of 15,000 individuals up to 125 times faster than existing tools, with identical results, and PCA of 150,000 individuals using flashpca completed in 4 hours. The increasing size of SNP datasets will make tools such as flashpca essential as traditional approaches will not adequately scale. This approach will also help to scale other applications that leverage PCA or eigen-decomposition to substantially larger datasets.

Genomics