Search bioRxivSearch

Biology subjects

Chen, C.-Y.

Publications and source records attributed to Chen, C.-Y..

28 records · Page 2Linked to original sources

Genetic validation of bipolar disorder identified by automated phenotyping using electronic health records

Bipolar disorder (BD) is a heritable mood disorder characterized by episodes of mania and depression. Although genomewide association studies (GWAS) have successfully identified genetic loci contributing to BD risk, sample size has become a rate-limiting obstacle to genetic discovery. Electronic health records (EHRs) represent a vast but relatively untapped resource for high-throughput phenotyping. As part of the International Cohort Collection for Bipolar Disorder (ICCBD), we previously validated automated EHR-based phenotyping algorithms for BD against in-person diagnostic interviews (Castro et al. 2015). Here, we establish the genetic validity of these phenotypes by determining their genetic correlation with traditionally-ascertained samples. Case and control algorithms were derived from structured and narrative text in the Partners Healthcare system comprising more than 4.6 million patients over 20 years. Genomewide genotype data for 3,330 BD cases and 3,952 controls of European ancestry were used to estimate SNP-based heritability (h2g) and genetic correlation(rg) between EHR-based phenotype definitions and traditionally-ascertained BD cases in GWAS by the ICCBD and Psychiatric Genomics Consortium (PGC) using LD score regression. We evaluated BD cases identified using 4 EHR-based algorithms: an NLP-based algorithm (95-NLP) and 3 rule-based algorithms using codified EHR with decreasing levels of stringency - \"coded-strict\", \"coded-broad\", and \"coded-broad based on a single clinical encounter\" (coded-broad-SV). The analytic sample comprised 862 95-NLP, 1,968 coded-strict, 2,581 coded-broad, 408 coded-broad-SV BD cases, and 3,952 controls. The estimated h2g were 0.24 (p=0.015), 0.09 (p=0.064), 0.13 (p=0.003), 0.00 (p=0.591) for 95-NLP, coded-strict, coded-broad and coded-broad-SV BD, respectively. The h2g for all EHR-based cases combined except coded-broad-SV (excluded due to 0 h2g) was 0.12 (p=0.004). These h2g were lower or similar to the h2g observed by the ICCBD+PGCBD (0.23, p=3.17E-80, total N=33,181). However, the rg between ICCBD+PGCBD and the EHR-based cases were high for 95-NLP (0.66, p=3.69x10-5), coded-strict (1.00, p=2.40x10-4), and coded-broad (0.74, p=8.11x10-7). The rg between EHR-based BDs ranged from 0.90 to 0.98. These results provide the first genetic validation of automated EHR-based phenotyping for BD and suggest that this approach identifies cases that are highly genetically correlated with those ascertained through conventional methods. High throughput phenotyping using the large data resources available in EHRs represents a viable method for accelerating psychiatric genetic research.

genetics

Spatial vision by macaque midbrain

Visual brain areas exhibit tuning characteristics that are well suited for image statistics present in our natural environment. However, visual sensation is an active process, and if there are any brain areas that ought to be particularly 'in tune' with natural scene statistics, it would be sensory-motor areas critical for guiding behavior. Here we found that the primate superior colliculus, a structure instrumental for rapid visual exploration with saccades, detects low spatial frequencies, which are the most prevalent in natural scenes, much more rapidly than high spatial frequencies. Importantly, this accelerated detection happens independently of whether a neuron is more or less sensitive to low spatial frequencies to begin with. At the population level, the superior colliculus additionally over-represents low spatial frequencies in neural response sensitivity, even at near-foveal eccentricities. Thus, the superior colliculus possesses both temporal and response gain mechanisms for efficient gaze realignment in low-spatial-frequency dominated natural environments.

neuroscience

Incomplete inhibition of HIV infection results in more HIV infected lymph node cells by reducing cell death

HIV has been reported to be cytotoxic in vitro and in lymph node infection models. Using a computational approach, we found that partial inhibition of transmission which involves multiple virions per cell could lead to increased numbers of live infected cells if the number of viral DNA copies remains above one after inhibition, as eliminating the surplus viral copies reduces cell death. Using a cell line, we observed increased numbers of live infected cells when infection was partially inhibited with the antiretroviral efavirenz or neutralizing antibody. We then used efavirenz at concentrations reported in lymph nodes to inhibit lymph node infection by partially resistant HIV mutants. We observed more live infected lymph node cells, but with fewer HIV DNA copies per cell, relative to no drug. Hence, counterintuitively, limited attenuation of HIV transmission per cell may increase live infected cell numbers in environments where the force of infection is high.

bioinformatics

Widespread pleiotropy confounds causal relationships between complex traits and diseases inferred from Mendelian randomization

A fundamental assumption in inferring causality of an exposure on complex disease using Mendelian randomization (MR) is that the genetic variant used as the instrumental variable cannot have pleiotropic effects. Violation of this no pleiotropy assumption can cause severe bias. Emerging evidence have supported a role for pleiotropy amongst disease-associated loci identified from GWA studies. However, the impact and extent of pleiotropy on MR is poorly understood. Here, we introduce a method called the Mendelian Randomization Pleiotropy RESidual Sum and Outlier (MR-PRESSO) test to detect and correct for pleiotropy in multi-instrument summary-level MR testing. We show using simulations that existing approaches are less sensitive to the detection of pleiotropy when it occurs in a subset of instrumental variables, as compared to MR-PRESSO. Next, we show that pleiotropy is widespread in MR, occurring in 41% amongst significant causal relationships (out of 4,250 MR tests total) from pairwise comparisons of 82 complex traits and diseases from summary level genome-wide association data. We demonstrate that pleiotropy causes distortion between-168% and 189% of the causal estimate in MR. Furthermore, pleiotropy induces false positive causal relationships-defined as those causal estimates that were no longer statistically significant in the pleiotropy corrected MR test but were previously significant in the naive MR test-in up to 10% of the MR tests using a P < 0.05 cutoff that is commonly used in MR studies. Finally, we show that MR-PRESSO can correct for distortion in the causal estimate in most cases. Our results demonstrate that pleiotropy is widespread and pervasive, and must be properly corrected for in order to maintain the validity of MR.

genomics

The Genomic Landscape Of Tree Rot In Phellinus noxius And Its Hymenochaetales Members

The order Hymenochaetales of white rot fungi contain some of the most aggressive wood decayers causing tree deaths around the world. Despite their ecological importance and the impact of diseases they cause, little is known about the evolution and transmission patterns of these pathogens. Here, we sequenced and undertook comparative genomics analyses of Hymenochaetales genomes using brown root rot fungus Phellinus noxius, wood-decomposing fungus Phellinus lamaensis, laminated root rot fungus Phellinus sulphurascens, and trunk pathogen Porodaedalea pini. Many gene families of lignin-degrading enzymes were identified from these fungi, reflecting their ability as white rot fungi. Comparing against distant fungi highlighted the expansion of 1,3-beta-glucan synthases in P. noxius, which may account for its fast-growing attribute. We identified 13 linkage groups conserved within Agaricomycetes, suggesting the evolution of stable karyotypes. We determined that P. noxius has a bipolar heterothallic mating system, with unusual highly expanded ~60 kb A locus as a result of accumulating gene transposition. We investigated the population genomics of 60 P. noxius isolates across multiple islands of the Asia Pacific region. Whole-genome sequencing showed this multinucleate species contains abundant poly-allelic single-nucleotide-polymorphisms (SNPs) with atypical allele frequencies. Different patterns of intra-isolate polymorphism reflect mono-/heterokaryotic states which are both prevalent in nature. We have shown two genetically separated lineages with one spanning across many islands despite the geographical barriers. Both populations possess extraordinary genetic diversity and show contrasting evolutionary scenarios. These results provide a framework to further investigate the genetic basis underlying the fitness and virulence of white rot fungi.

genomics

Understanding Tissue-specific Gene Regulation

Although all human tissues carry out common processes, tissues are distinguished by gene expres-sion patterns, implying that distinct regulatory programs control tissue-specificity. In this study, we investigate gene expression and regulation across 38 tissues profiled in the Genotype-Tissue Expression project. We find that network edges (transcription factor to target gene connections) have higher tissue-specificity than network nodes (genes) and that regulating nodes (transcription factors) are less likely to be expressed in a tissue-specific manner as compared to their targets (genes). Gene set enrichment analysis of network targeting also indicates that regulation of tissue-specific function is largely independent of transcription factor expression. In addition, tissue-specific genes are not highly targeted in their corresponding tissue-network. However, they do assume bottleneck positions due to variability in transcription factor targeting and the influence of non-canonical regulatory interactions. These results suggest that tissue-specificity is driven by context-dependent regulatory paths, providing transcriptional control of tissue-specific processes.

genomics

A network-based approach to eQTL interpretation and SNP functional characterization

Expression quantitative trait locus (eQTL) analysis associates genotype with gene expression, but most eQTL studies only include cis-acting variants and generally examine a single tissue. We used data from 13 tissues obtained by the Genotype-Tissue Expression (GTEx) project v6.0 and, in each tissue, identified both cis- and trans-eQTLs. For each tissue, we represented significant associations between single nucleotide polymorphisms (SNPs) and genes as edges in a bipartite network. These networks are organized into dense, highly modular communities often representing coherent biological processes. Global network hubs are enriched in distal gene regulatory regions such as enhancers, but are devoid of disease-associated SNPs from genome wide association studies. In contrast, local, community-specific network hubs (core SNPs) are preferentially located in regulatory regions such as promoters and enhancers and highly enriched for trait and disease associations. These results provide help explain how many weak-effect SNPs might together influence cellular function and phenotype.

genomics

Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse data

Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. We find analysis of large RNA-Seq data sets requires both careful quality control and that one account for sparsity due to the heterogeneity intrinsic in multi-group studies. An R package instantiating our method for large-scale RNA-Seq normalization and preprocessing, YARN, is available at bioconductor.org/packages/yarn.\n\nHighlightsO_LIOverview of assumptions used in preprocessing and normalization\nC_LIO_LIPipeline for preprocessing, quality control, and normalization of large heterogeneous data\nC_LIO_LIA Bioconductor package for the YARN pipeline and easy manipulation of count data\nC_LIO_LIPreprocessed GTEx data set using the YARN pipeline available as a resource\nC_LI

bioinformatics

Transcriptional landscape of cell lines and their tissues of origin

Cell lines are an indispensable tool in biomedical research and often used as surrogates for tissues. An important question is how well a cell lines transcriptional and regulatory processes reflect those of its tissue of origin. We analyzed RNA-Seq data from GTEx for 127 paired Epstein-Barr virus transformed lymphoblastoid cell lines and whole blood samples; and 244 paired fibroblast cell lines and skin biopsies. A combination of gene expression and network analyses shows that while cell lines carry the expression signatures of their primary tissues, albeit at reduced levels, they also exhibit changes in their patterns of transcription factor regulation. Cell cycle genes are over-expressed in cell lines compared to primary tissue, and they have a reduction of repressive transcription factor targeting. Our results provide insight into the expression and regulatory alterations observed in cell lines and suggest that these changes should be considered when using cell lines as models.\n\nHighlightsO_LICell lines differ from their source tissues in gene expression and regulation\nC_LIO_LIDistinct cell lines share altered patterns of cell cycle regulation\nC_LIO_LICell cycle genes are less strongly targeted by repressive TFs in cell lines\nC_LIO_LICell lines share expression with their source tissue, but at reduced levels\nC_LI

genomics

Sexual dimorphism in gene expression and regulatory networks across human tissues

Sexual dimorphism manifests in many diseases and may drive sex-specific therapeutic responses. To understand the molecular basis of sexual dimorphism, we conducted a comprehensive assessment of gene expression and regulatory network modeling in 31 tissues using 8716 human transcriptomes from GTEx. We observed sexually dimorphic patterns of gene expression involving as many as 60% of autosomal genes, depending on the tissue. Interestingly, sex hormone receptors do not exhibit sexually dimorphic expression in most tissues; however, differential network targeting by hormone receptors and other transcription factors (TFs) captures their downstream sexually dimorphic gene expression. Furthermore, differential network wiring was found extensively in several tissues, particularly in brain, in which not all regions exhibit strong differential expression. This systems-based analysis provides a new perspective on the drivers of sexual dimorphism, one in which a repertoire of TFs plays important roles in sex-specific rewiring of gene regulatory networks.\n\nHighlightsO_LISexual dimorphism manifests in both gene expression and gene regulatory networks\nC_LIO_LISubstantial sexual dimorphism in regulatory networks was found in several tissues\nC_LIO_LIMany differentially regulated genes are not differentially expressed\nC_LIO_LISex hormone receptors do not exhibit sexually dimorphic expression in most tissues\nC_LI

genomics