Search bioRxivSearch

Biology subjects

Choi, K.

Publications and source records attributed to Choi, K..

14 recordsLinked to original sources

Accurate estimation of single cell allele-specific gene expression using all reads and combining information across cells

Allele-specific expression (ASE) at single-cell resolution is a critical tool for understanding the stochastic and dynamic features of gene expression. However, low read coverage and high biological variability present challenges for analyzing ASE. We propose a new method for ASE analysis from single cell RNA-Seq data that accurately classifies allelic expression states and improves estimation of allelic proportions by pooling information across cells.

bioinformatics

Continuous single cell transcriptome dynamics from pluripotency to hemangiogenic lineage

Blood and endothelial cells arise from hemangiogenic progenitors that are specified from FLK1-expressing mesoderm by the transcription factor ETV2. FLK1 mesoderm also contributes to other tissues, including vascular smooth muscle (VSM) and cardiomyocytes. However, the developmental process of FLK1 mesoderm generation and its allocation to various cell fates remain obscure. Recent single cell RNA-sequencing (scRNA-seq) studies of early stages of embryos, or in vitro differentiated human embryonic stem (ES) cells have provided unprecedented information on the spatiotemporal resolution of cells in embryogenesis. These snapshots nonetheless offer insufficient information on dynamic developmental processes due to inadvertently missing intermediate states and unavoidable batch effects. Here we performed scRNA-seq of in vitro differentiated ES cells as well as extraembryonic yolk sac cells, which contain the very first arising hemangiogenic and VSM lineages, to capture the continuous developmental process leading to hemangiogenesis. We found that hemangiogenic progenitors from ES cells develop through intermediate gastrulation stages, which were gradually specified by relay-like highly overlapping transcription factor modules. Unexpectedly, VSM and hemangiogenic lineages share the closest transcriptional program. Moreover, transcriptional program of the Flk1 mesoderm was maintained in the VSM lineage, suggesting the VSM lineage may be the default pathway of FLK1 mesoderm. We also identified cell adhesion signals possibly contributing to ETV2-mediated activation of the hemangiogenic program. This continuous transcriptome map will facilitate both basic and applied studies of mesoderm and its derivatives.

developmental biology

Testing Causal Bidirectional Influences between Physical Activity and Depression using Mendelian Randomization

BackgroundBurgeoning evidence from randomized controlled trials and prospective cohort studies suggests that physical activity protects against depression, pointing to a potential modifiable target for prevention. However, the direction of this inverse association is not clear: physical activity may reduce risk for depression, and/or depression may result in decreased physical activity. Here, we used bidirectional two-sample Mendelian randomization (MR) to test causal influences between physical activity and depression.\n\nMethodsFor genetic instruments, we selected independent top SNPs associated with major depressive disorder (MDD, N = 143,265) and two physical activity phenotypes--self-reported (N = 377,234) and objective accelerometer-based (N = 91,084)--from the largest available, non-overlapping genome-wide association results. We used two sets of genetic instruments: (1) only SNPs previously reported as genome-wide significant, and (2) top SNPs meeting a more relaxed threshold (p < 1x10-7). For each direction of influence, we combined the MR effect estimates from each instrument SNP using inverse variance weighted (IVW) meta-analysis, along with other standard MR methods such as weighted median, MR-Egger, and MR-PRESSO.\n\nResultsWe found evidence for protective influences of accelerometer-based activity on MDD (IVW odds ratio (OR) = 0.74 for MDD per 1 SD unit increase in average acceleration, 95% confidence interval (CI) = 0.59-0.92, p =.006) when using SNPs meeting the relaxed threshold (i.e., 10 versus only 2 genome-wide significant SNPs, which provided insufficient data for sensitivity analyses). In contrast, we found no evidence for negative influences of MDD on accelerometer-based activity (IVW b = 0.04 change in average acceleration for MDD versus control status, 95% CI = -0.43-0.51, p =.87). Furthermore, we did not see evidence for causal influences between self-reported activity and MDD, in either direction and regardless of instrument SNP criteria.\n\nDiscussionWe apply MR for the first time to examine causal influences between physical activity and MDD. We discover that objectively measured--but not self-reported--physical activity is inversely associated with MDD. Of note, prior work has shown that accelerometer-based physical activity is more heritable than self-reported activity, in addition to being more representative of actual movement. Our findings validate physical activity as a protective factor for MDD and point to the importance of objective measurement of physical activity in epidemiological studies in relation to mental health. Overall, this study supports the hypothesis that enhancing physical activity is an effective prevention strategy for depression.

genetics

Prospective Study of Polygenic Risk, Protective Factors, and Incident Depression Following Combat Deployment in US Army Soldiers

BackgroundWhereas genetic susceptibility increases risk for major depressive disorder (MDD), non-genetic protective factors may mitigate this risk. In a large-scale prospective study of US Army soldiers, we examined whether trait resilience and/or unit cohesion could buffer against the onset of MDD following combat deployment, even in soldiers at high polygenic risk.\n\nMethodsData were analyzed from 4,182 soldiers of European ancestry assessed before and after their deployment to Afghanistan. Incident MDD was defined as no MDD episode at predeployment, followed by a MDD episode following deployment. Polygenic risk scores were constructed from the largest available MDD genome-wide association study. We first examined main effects of the MDD PRS and each protective factor on incident MDD. We then tested effects of each protective factor on incident MDD across strata of polygenic risk.\n\nResultsPolygenic risk showed a dose-response relationship to depression, such that soldiers at high polygenic risk had greatest odds for incident MDD. Both unit cohesion and trait resilience were prospectively associated with reduced risk for incident MDD. Notably, the protective effect of unit cohesion persisted even in soldiers at highest polygenic risk.\n\nConclusionsPolygenic risk was associated with new-onset MDD in deployed soldiers. However, unit cohesion--an index of perceived support and morale--was protective against incident MDD even among those at highest genetic risk, and may represent a potent target for promoting resilience in vulnerable soldiers. Findings illustrate the value of combining genomic and environmental data in a prospective design to identify robust protective factors for mental health.

genetics

Inferring Reaction Networks using Perturbation Data

In this paper we examine the use of perturbation data to infer the underlying mechanistic dynamic model. The approach uses an evolutionary strategy to evolve networks based on a fitness criterion that measures the difference between the experimentally determined set of perturbation data and proposed mechanistic models. At present we only deal with reaction networks that use mass-action kinetics employing uni-uni, bi-uni, uni-bi and bi-bi reactions. The key to our approach is to split the algorithm into two phases. The first phase focuses on evolving network topologies that are consistent with the perturbation data followed by a second phase that evolves the parameter values. This results in almost an exact match between the evolved network and the original network from which the perturbation data was generated from. We test the approach on four models that include linear chain, feed-forward loop, cyclic pathway and a branched pathway. Currently the algorithm is implemented using Python and libRoadRunner but could at a later date be rewritten in a compiled language to improve performance. Future studies will focus on the impact of noise in the perturbation data on convergence and variability in the evolved parameter values and topologies. In addition we will investigate the effect of nonlinear rate laws on generating unique solutions.

systems biology

Tissue-specific trans regulation of the mouse epigenome

Although a variety of writers, readers, and erasers of epigenetic modifications are known, we have little information about the underlying regulatory systems controlling the establishment and maintenance of the epigenetic landscape, which varies greatly among cell types. Here, we have explored how natural genetic variation impacts the epigenome in mice. Studying levels of H3K4me3, a histone modification at sites such as promoters, enhancers, and recombination hotspots, we found tissue-specific trans-regulation of H3K4me3 levels in four highly diverse cell types: male germ cells, embryonic stem (ES) cells, hepatocytes and cardiomyocytes. To identify the genetic loci involved, we measured H3K4me3 levels in male germ cells in a mapping population of 60 BXD recombinant inbred lines, identifying extensive trans-regulation primarily controlled by six major histone quantitative trait loci (hQTL). These chromatin regulatory loci act dominantly to suppress H3K4me3, which at hotspots reduces the likelihood of subsequent DNA double-strand breaks. QTL locations do not correspond with enzyme known to metabolize chromatin features. Instead their locations match clusters of zinc finger genes, making these possible candidates that explain the dominant suppression of H3K4me3. Collectively, these data describe an extensive, tissue-specific set of chromatin regulatory loci that control functionally related chromatin sites.

genomics

Interhomolog polymorphism shapes meiotic crossover within RAC1 and RPP13 disease resistance genes

During meiosis chromosomes undergo DNA double-strand breaks (DSBs), which can produce crossovers via interhomolog repair. Meiotic recombination frequency is variable along chromosomes and concentrates in narrow hotspots. We mapped crossovers within Arabidopsis thaliana hotspots located within the RAC1 and RPP13 disease resistance genes, using varying haplotypic combinations. We observed a negative non-linear relationship between interhomolog divergence and crossover frequency, consistent with polymorphism suppressing crossover repair of DSBs. Anti-recombinase mutants fancm, recq4a recq4b, figl1 and msh2, or lines with increased HEI10 dosage, are known to show increased crossovers. Surprisingly, RAC1 crossovers were either unchanged or decreased in these genetic backgrounds. We employed deep-sequencing of crossovers to examine recombination topology within RAC1, in wild type, fancm and recq4a recq4b mutant backgrounds. The RAC1 recombination landscape was broadly conserved in anti-recombinase mutants and showed a negative relationship with interhomolog divergence. However, crossovers at the RAC1 5-end were relatively suppressed in recq4a recq4b backgrounds, indicating that local context influences recombination outcomes. Our results demonstrate the importance of interhomolog divergence in shaping recombination within plant disease resistance genes and crossover hotspots.

genetics

Tellurium Notebooks - An Environment for Dynamical Model Development, Reproducibility, and Reuse

The considerable difficulty encountered in reproducing the results of published dynamical models limits validation, exploration and reuse of this increasingly large biomedical research resource. To address this problem, we have developed Tellurium Notebook, a software system that facilitates building reproducible dynamical models and reusing models by 1) supporting the COMBINE archive format during model development for capturing model information in an exchangeable format and 2) enabling users to easily simulate and edit public COMBINE-compliant models from public repositories to facilitate studying model dynamics, variants and test cases. Tellurium Notebook, a Python-based Jupyter-like environment, is designed to seamlessly inter-operate with these community standards by automating conversion between COMBINE standards formulations and corresponding in-line, human-readable representations. Thus, Tellurium brings to systems biology the strategy used by other literate notebook systems such as Mathematica. These capabilities allow users to edit every aspect of the standards-compliant models and simulations, run the simulations in-line, and re-export to standard formats. We provide several use cases illustrating the advantages of our approach and how it allows development and reuse of models without requiring technical knowledge of standards. Adoption of Tellurium should accelerate model development, reproducibility and reuse.\n\nAuthor summaryThere is considerable value to systems and synthetic biology in creating reproducible models. An essential element of reproducibility is the use of community standards, an often challenging undertaking for modelers. This article describes Tellurium Notebook, a tool for developing dynamical models that provides an intuitive approach to building and reusing models built with community standards. Tellurium automates embedding human-readable representations of COMBINE archives in literate coding notebooks, bringing to systems biology this strategy central to other literate notebook systems such as Mathematica. We show that the ability to easily edit this human-readable representation enables users to test models under a variety of conditions, thereby providing a way to create, reuse, and modify standard-encoded models and simulations, regardless of the users level of technical knowledge of said standards.

systems biology

Single-Cell Immune Map of Breast Carcinoma Reveals Diverse Phenotypic States Driven by the Tumor Microenvironment

Knowledge of immune cell phenotypes in the tumor microenvironment is essential for understanding mechanisms of cancer progression and immunotherapy response. We created an immune map of breast cancer using single-cell RNA-seq data from 45,000 immune cells from eight breast carcinomas, as well as matched normal breast tissue, blood, and lymph node. We developed a preprocessing pipeline, SEQC, and a Bayesian clustering and normalization method, Biscuit, to address computational challenges inherent to single-cell data. Despite significant similarity between normal and tumor tissue-resident immune cells, we observed continuous tumor-specific phenotypic expansions driven by environmental cues. Analysis of paired single-cell RNA and T cell receptor (TCR) sequencing data from 27,000 additional T cells revealed the combinatorial impact of TCR utilization on phenotypic diversity. Our results support a model of continuous activation in T cells and do not comport with the macrophage polarization model in cancer, with important implications for characterizing tumor-infiltrating immune cells.

immunology

Context-Dependent Genetic Regulation

Cells process extra-cellular signals with multiple layers of complex biological networks. Due to the stochastic nature of the networks, the signals become significantly noisy within the cells and in addition, due to the nonlinear nature of the networks, the signals become distorted, shifted, and (de-)amplified. Such nonlinear signal processing can lead to non-trivial cellular phenotypes such as cell cycles, differentiation, cell-to-cell communication, and homeostasis. These nonlinear pheno-types, when observed at the cell population levels, can be quite different from the single-cell level observation. As one of the underlying mechanisms behind this difference, we report the interplay between nonlinearity and stochasticity in genetic regulation. Here we show that nonlinear genetic regulation, characterized at the cellular population level, can be affected by cell-to-cell variability in the regulatory factor concentrations. The observed genetic regulation at the cell population is shown to be significantly dependent on the upstream DNA sequences of the regulator, in particular, 5 untranslated region. This indicates that genetic regulation observed at the cell population level can be significantly dependent on its genetic context, and that its characterization needs a careful attention on noise propagation.\n\nOne Sentence SummaryGenetic regulation observed at the cell population level can be significantly affected by cell-to-cell variability in the regulatory factor copy numbers, indicating that the observed regulation is dependent on 5 UTR of the regulator coding gene.

synthetic biology

Hierarchical Analysis of Multi-mapping RNA-Seq Reads Improves the Accuracy of Allele-specific Expression

Allele-specific expression (ASE) refers to the differential abundance of the allelic copies of a transcript. Direct RNA sequencing (RNA-Seq) can provide quantitative estimates of ASE for genes with transcribed polymorphisms. However, estimating ASE is challenging due to ambiguities in read alignment. Current approaches do not account for the hierarchy of multiple read alignments to genes, isoforms, and alleles. We have developed EMASE (Expectation-Maximization for Allele Specific Expression), an integrated approach to estimate total gene expression, ASE, and isoform usage based on hierarchical allocation of multi-mapping reads. In simulations, EMASE outperforms standard ASE estimation methods. We apply EMASE to RNA-Seq data from F1 hybrid mice where we observe widespread ASE associated with cis-acting polymorphisms and a small number of parent-of-origin effects at known imprinted genes. The EMASE software is freely available under GNU license at https://github.com/churchill-lab/emase and it can be adapted to other sequencing applications.

bioinformatics

Nucleosomes and DNA methylation shape meiotic DSB frequency in Arabidopsis transposons and gene regulatory regions

Meiotic recombination initiates via DNA double strand breaks (DSBs) generated by SPO11 topoisomerase-like complexes. Recombination frequency varies extensively along eukaryotic chromosomes, with hotspots controlled by chromatin and DNA sequence. To map meiotic DSBs throughout a plant genome, we purified and sequenced Arabidopsis SPO11-1-oligonucleotides. DSB hotspots occurred in gene promoters, terminators and introns, driven by AT-sequence richness, which excludes nucleosomes and allows SPO11-1 access. A strong positive relationship was observed between SPO11-1 DSBs and final crossover levels. Euchromatic marks promote recombination in fungi and mammals, and consistently we observe H3K4me3 enrichment in proximity to DSB hotspots at gene 5-ends. Repetitive transposons are thought to be recombination-silenced during meiosis, in order to prevent non-allelic interactions and genome instability. Unexpectedly, we found strong DSB hotspots in nucleosome-depleted Helitron/Pogo/Tc1/Mariner DNA transposons, whereas retrotransposons were coldspots. Hotspot transposons are enriched within gene regulatory regions and in proximity to immunity genes, suggesting a role as recombination-enhancers. As transposon mobility in plant genomes is restricted by DNA methylation, we used the met1 DNA methyltransferase mutant to investigate the role of heterochromatin on the DSB landscape. Epigenetic activation of transposon meiotic DSBs occurred in met1 mutants, coincident with reduced nucleosome occupancy, gain of transcription and H3K4me3. Increased met1 SPO11-1 DSBs occurred most strongly within centromeres and Gypsy and CACTA/EnSpm coldspot transposons. Together, our work reveals complex interactions between chromatin and meiotic DSBs within genes and transposons, with significance for the diversity and evolution of plant genomes.

genomics

Epigenetic activation of meiotic recombination in Arabidopsis centromeres via loss of H3K9me2 and non-CG DNA methylation

Eukaryotic centromeres contain the kinetochore, which connects chromosomes to the spindle allowing segregation. During meiosis centromeres are suppressed for crossovers, as recombination in these regions can cause chromosome mis-segregation. Plant centromeres are surrounded by repetitive, transposon-dense heterochromatin that is epigenetically silenced by histone 3 lysine 9 dimethylation (H3K9me2), and DNA methylation in CG and non-CG sequence contexts. Here we show that disruption of Arabidopsis H3K9me2 and non-CG DNA methylation pathways increases meiotic DNA double strand breaks (DSBs) within centromeres, whereas crossovers increase within pericentromeric heterochromatin. Increased pericentromeric crossovers in H3K9me2/non-CG mutants occurs in both inbred and hybrid backgrounds, and involves the interfering crossover repair pathway. Epigenetic activation of recombination may also account for the curious tendency of maize transposon Ds to disrupt CHROMOMETHYLASE3 when launched from proximal loci. Thus H3K9me2 and non-CG DNA methylation exert differential control of meiotic DSB and crossover formation in centromeric and pericentromeric heterochromatin.

genomics

The Effects of Sex and Diet on Physiology and Liver Gene Expression in Diversity Outbred Mice

Inter-individual variation in metabolic health and adiposity is driven by many factors. Diet composition and genetic background and the interactions between these two factors affect adiposity and related traits such as circulating cholesterol levels. In this study, we fed 850 Diversity Outbred mice, half females and half males, with either a standard chow diet or a high fat, high sucrose diet beginning at weaning and aged them to 26 weeks. We measured clinical chemistry and body composition at early and late time points during the study, and liver transcription at euthanasia. Males weighed more than females and mice on a high fat diet generally weighed more than those on chow. Many traits showed sex- or diet-specific changes as well as more complex sex by diet interactions. We mapped both the physiological and molecular traits and found that the genetic architecture of the physiological traits is complex, with many single locus associations potentially being driven by more than one polymorphism. For liver transcription, we find that local polymorphisms affect constitutive and sex-specific transcription, but that the response to diet is not affected by local polymorphisms. We identified two loci for circulating cholesterol levels. We performed mediation analysis by mapping the physiological traits, given liver transcript abundance and propose several genes that may be modifiers of the physiological traits. By including both physiological and molecular traits in our analyses, we have created deeper phenotypic profiles to identify additional significant contributors to complex metabolic outcomes such as polygenic obesity. We make the phenotype, liver transcript and genotype data publicly available as a resource for the research community.

genetics