Search bioRxivSearch

Biology subjects

Choi, H.

Publications and source records attributed to Choi, H..

12 recordsLinked to original sources

AnnFlux: object-conditioned neural stochastic differential equations for single-cell perturbation dynamics

Single-cell perturbation profiling measures responses to genetic and chemical interventions, yet most models learn a static map, ignoring how populations move over time and how perturbations combine. AnnFlux, an object-conditioned stochastic differential equation, learns a drift field in latent cell-state space. Conditioning on the perturbing object makes the field queryable one object at a time, yielding per-object drifts comparable across genes and drugs. By learning a drift field tailored to each perturbation context, it interpolates a held-out timepoint in an epithelial-mesenchymal transition time course and predicts unseen perturbations. Beyond point estimates, AnnFlux improves distributional fidelity and predicts responses to held-out perturbation combinations. An IFN-response signature predicted by AnnFlux was associated with TLS proximity in an independent pan-cancer spatial atlas. This framework maps perturbation-driven cell-state evolution as continuous trajectories and represents unseen perturbations using prior-knowledge embeddings.

bioinformatics

Moving beyond P values: Everyday data analysis with estimation plots

Introduction Introduction The two-groups design is... Significance testing obscures... Even when fully visualized,... Five key advantages of... Estimation graphics are... Conclusion Author Contributions Funding sources Code Availability Guide to using DABEST Web application Google Colaboratory Python Matlab DABEST-Python in R References Over the past 75 years, a number of statisticians have advised that the data-analysis method known as null- hypothesis significance testing (NHST) should be deprecated (Berkson, 1942; Halsey et al., 2015). The l ...

bioinformatics

iOmicsPASS: a novel method for integration of multi-omics data over biological networks and discovery of predictive subnetworks

We developed iOmicsPASS, an intuitive method for network-based multi-omics data integration and detection of biological subnetworks for phenotype prediction. The method converts abundance measurements into co-expression scores of biological networks and uses a powerful phenotype prediction method adapted for network-wise analysis. Simulation studies show that the proposed data integration approach considerably improves the quality of predictions. We illustrate iOmicsPASS through the integration of quantitative multi-omics data using transcription factor regulatory network and protein-protein interaction network for cancer subtype prediction. Our analysis of breast cancer data identifies network signatures surrounding established markers of molecular subtypes. The analysis of colorectal cancer data highlights a protein interactome surrounding key proto-oncogenes as predictive features of subtypes, rendering them more biologically interpretable than the approaches integrating data without a priori relational information. However, the results indicate that current molecular subtyping is overly dependent on transcriptomic data and crude integrative analysis fails to account for molecular heterogeneity in other -omics data. The analysis also suggest that tumor subtypes are not mutually exclusive and future subtyping should therefore consider multiplicity in assignments.\n\nAvailability: https://github.com/cssblab/iOmicsPASS

systems biology

Addition of Degenerate Bases to DNA-based Data Storage for Increased Information Capacity

Introductory paragraphDNA-based data storage has emerged as a promising method to satisfy the exponentially increasing demand for information storage. However, practical implementation of DNA-based data storage remains a challenge because of the high cost of DNA per unit data. Here, we propose the use of eleven degenerate bases as encoding characters in addition to A, C, G, and T, which increases the information capacity (the amount of data that can be stored per length of DNA sequence designed) and reduce the cost of DNA per unit data. Using the proposed method, we experimentally achieved an information capacity of 3.37 bits/character, which is more than twice when compared to the highest information capacity previously achieved. Finally, the platform was projected to reduce the cost of DNA-based data storage by 50%.

synthetic biology

mTOR-dependent phosphorylation controls TFEB nuclear export

The transcriptional activation of catabolic processes during starvation is induced by the nuclear translocation and consequent activation of transcription factor EB (TFEB), a master modulator of autophagy and lysosomal biogenesis. However, how TFEB is inactivated upon nutrient re-feeding is currently unknown. Here we show that TFEB subcellular localization is dynamically controlled by its continuous shuttling between the cytosol and the nucleus, with the nuclear export representing a limiting step. TFEB nuclear export is mediated by CRM1 and is modulated by nutrient availability via mTOR-dependent hierarchical multisite phosphorylation of serines S142 and S138, which are localized in proximity of a nuclear export signal (NES). Our data reveal that modulation of TFEB nuclear export via phosphorylation plays a major role in the modulation of TFEB localization and activity.

cell biology

Synchronization dependent on spatial structures of a mesoscopic whole-brain network

We study how the spatial structure of connectivity shapes synchronization in a system of coupled phase oscillators on a mammalian whole-brain network at the mesoscopic level. Complex structural connectivity of the mammalian brain is believed to underlie the versatility of neural computations. The Allen Mouse Brain Connectivity Atlas constructed from viral tracing experiments together with a new mapping algorithm reveals that the connectivity has a significant spatial dependence: the connection strength decreases with distance between the regions, following a power law. However, there are a number of residuals above the power-law fit, predominantly for long-range connections. We show how these strong connections between distal brain regions promote rapid transitions between highly localized synchronization and more global synchronization as the amount of dispersion in the frequency distribution changes. This may explain the brains ability to switch rapidly between global and modularized computations.

neuroscience

Quantifying multi-layered expression regulation in response to stress of the endoplasmic reticulum

The mammalian response to endoplasmic reticulum (ER) stress dynamically affects all layers of gene expression regulation. We quantified transcript and protein abundance along with footprints of ribosomes and non-ribosomal proteins for thousands of genes in cervical cancer cells responding to treatment with tunicamycin or hydrogen peroxide over an eight hour time course. We identify shared and stress-specific significant regulatory events at the transcriptional and post-transcriptional level and at different phases of the experiment. ER stress regulators increase transcription and translation at different times supporting an adaptive response. ER stress also induces translation of genes from serine biosynthesis and one-carbon metabolism indicating a shift in energy production. Discordant regulation of DNA repair genes suggests transcriptional priming in which delayed translation fine-tunes the early change in the transcriptome. Finally, case studies on stress-dependent alternative splicing and protein-mRNA binding demonstrate the ability of this resource to generate hypotheses for new regulatory mechanisms.

systems biology

PTMscape: an open source tool to predict generic post-translational modifications and map hotspots of modification crosstalk

While tandem mass spectrometry can now detect post-translational modifications (PTM) at the proteome scale, reported modification sites are often incomplete and include false positives. Computational approaches can complement these datasets by additional predictions, but most available tools are tailored for single modifications and each tool uses different features for prediction. We developed an R package called PTMscape which predicts modifications sites across the proteome based on a unified and comprehensive set of descriptors of the physico-chemical microenvironment of modified sites, with additional downstream analysis modules to test enrichment of individual or pairs of modifications in functional protein regions. PTMscape is generic in the ability to process any major modifications, such as phosphorylation and ubiquitination, while achieving the sensitivity and specificity comparable to single-PTM methods and outperforming other multi-PTM tools. Maintaining generalizability of the framework, we expanded proteome-wide coverage of five major modifications affecting different residues by prediction and performed combinatorial analysis for spatial co-occurrence of pairs of those modifications. This analysis revealed potential modification hotspots and crosstalk among multiple PTMs in key protein domains such as histone, protein kinase, and RNA recognition motifs, spanning various biological processes such as RNA processing, DNA damage response, signal transduction, and regulation of cell cycle. These results provide a proteome-scale analysis of crosstalk among major PTMs and can be easily extended to other modifications.\n\nContactall correspondence should be addressed to hwchoi@nus.edu.sg.

bioinformatics

Predicting aging of brain metabolic topography using variational autoencoder

Predicting future brain topography can give insight into neural correlates of aging and neurodegeneration. Due to variability in aging process, it has been challenging to precisely estimate brain topographical change according to aging. Here, we predict age-related brain metabolic change by generating future brain 18F-Fluorodeoxyglucose PET. A cross-sectional PET dataset of cognitively normal subjects with different age was used to develop a generative model. The model generated PET images using age information and characteristic individual features. Predicted regional metabolic changes were correlated with the real changes obtained by follow-up data. This model was applied to produce a brain metabolism aging movie by generating PET at different ages. Normal population distribution of brain metabolic topography at each age was estimated as well. In addition, a generative model using APOE4 status as well as age as inputs revealed a significant effect of APOE4 status on age-related metabolic changes particularly in the calcarine, lingual cortex, hippocampus and amygdala. It suggested APOE4 could be a factor affecting individual variability in age-related metabolic degeneration in normal elderly. This predictive model may not only be extended to understanding cognitive aging process, but apply to development of a preclinical biomarker for various brain disorders.

neuroscience

A risk stratification model for lung cancer based on gene coexpression network

Risk stratification model for lung cancer with gene expression profile is of great interest. Instead of the previously reported models based on individual prognostic genes, we aimed to develop a novel system-level risk stratification model for lung adenocarcinoma based on gene coexpression network. Using multiple microarray datasets obtained from lung adenocarcinoma, gene coexpression network analysis was performed to identify survival-related network modules. Representative genes of these network modules were selected and then, risk stratification model was constructed exploiting deep learning algorithm. The model was validated in two independent test cohorts. Survival analysis using univariate and multivariate Cox regression was performed using the output of the model to evaluate whether the model could predict patients overall survival independent of clinicopathological variables. Five network modules were significantly associated with patients survival. Considering prognostic significance and representativeness, genes of the two survival-related modules were selected for input data of the risk stratification model. The output of the model was significantly associated with patients overall survival in the two independent test sets as well as training set (p < 0.00001, p < 0.0001 and p = 0.02 for training set, test set 1 and 2, respectively). In multivariate analyses, the model was associated with patients prognosis independent of other clinical and pathological features. Our study presents a new perspective on incorporating gene coexpression networks into the gene expression signature, and the clinical application of deep learning in genomic data science for prognosis prediction.

bioinformatics

3D Mapping Reveals Network-specific Amyloid Progression and Subcortical Susceptibility.

Alzheimers disease is a progressive, neurodegenerative condition for which there is no cure. Prominent hypotheses posit that accumulation of beta-amyloid (A{beta}) peptides drives the neurodegeneration that underlies memory loss, however the spatial origins of the lesions remain elusive. Using SWITCH, we created a spatiotemporal map of A{beta} deposition in a mouse model of amyloidosis. We report that structures connected by the fornix show primary susceptibility to A{beta} accumulation and demonstrate that aggregates develop in increasingly complex networks with age. Notably, the densest early A{beta} aggregates occur in the mammillary body coincident with electrophysiological alterations. In later stages, the fornix itself also develops overt A{beta} burden. Finally, we confirm A{beta} in the mammillary body of postmortem patient specimens. Together, our data suggest that subcortical memory structures are particularly vulnerable to A{beta} deposition and that functional alterations within and physical propagation from these regions may underlie the affliction of increasingly complex networks.\n\nAuthor ContributionsRGC, KC, L-HT, ID conceived of the work and planned the experiments.\n\nRGC, HC, JW, LAW, CGY, FA, SMB performed experiments and analyzed data.\n\nHC built the custom microscope.\n\nRGC, L-HT, KC, ID wrote the manuscript.

neuroscience

A comprehensive survey of genetic variation in 20,691 subjects from four large cohorts

The Nurses Health Study (NHS), Nurses Health Study II (NHSII), Health Professionals Follow Up Study (HPFS) and the Physicians Health Study (PHS) have collected detailed longitudinal data on multiple exposures and traits for approximately 310,000 study participants over the last 35 years. Over 160,000 study participants across the cohorts have donated a DNA sample and to date, 20,691 subjects have been genotyped as part of genome-wide association studies (GWAS) of twelve primary outcomes. However, these studies utilized six different GWAS arrays making it difficult to conduct analyses of secondary phenotypes or share controls across studies. To allow for secondary analyses of these data, we have created three new datasets merged by platform family and performed imputation using a common reference panel, the 1,000 Genomes Phase I release. Here, we describe the methodology behind the data merging and imputation and present imputation quality statistics and association results from two GWAS of secondary phenotypes (body mass index (BMI) and venous thromboembolism (VTE)).\n\nWe observed the strongest BMI association for the FTO SNP rs55872725 ({beta}=0.45, p=3.48x10-22), and using a significance level of p=0.05, we replicated 19 out of 32 known BMI SNPs. For VTE, we observed the strongest association for the rs2040445 SNP (OR=2.17, 95% CI: 1.79-2.63, p=2.70x10-15), located downstream of F5 and also observed significant associations for the known ABO and F11 regions. This pooled resource can be used to maximize power in GWAS of phenotypes collected across the cohorts and for studying gene-environment interactions as well as rare phenotypes and genotypes.

epidemiology