Search bioRxiv⌕ Search

Biology subjects

Shin, E.-S.

Publications and source records attributed to Shin, E.-S..

4 recordsLinked to original sources

10,239 whole genomes with multiomic and clinical health information as the Korean population multiomic reference dataset

We present Korea10K, the largest genomic dataset of the Korean population, comprising 10,239 high-coverage whole genomes (mean depth 30x) with matched multiomic profiles and phenotype data. Korea10K achieves complete and near-complete discovery of very rare and ultra-rare alleles, respectively, at 9,000 Korean genomes. This dataset provides the high-quality population-specific imputation panel, enabling accurate inference of low-frequency variants. Admixture analyses confirm the genetic homogeneity of the Korean population, despite its diverse Y-chromosomal, mitochondrial, and HLA repertoires. This pattern reflects a long and continuous lineage history characterized by persistent internal admixture and genomic homogenization over thousands of years on the Korean peninsula. We also identified 16.8 million genomic variants that directly modify CG sites by creating or abolishing CG dinucleotides, providing the population-scale evidence of coordinated genomic-epigenomic regulatory mechanism in Koreans.

genomics↗

Identification of reactive CpGs and RNA expression in early COVID-19 through cis-eQTM analysis reflecting disease severity and recovery

Multi-omics analyses of severe COVID-19 cases are crucial in deciphering the complex interplay between genetic and epigenetic factors. Here, we present an analysis of Expression Quantitative Trait Methylation (eQTM) to investigate the complex interplay of methylation and gene expression pattern during the acute phase of severe COVID-19. We identified 16 differentially expressed genes and 30 nearby differentially methylated CpG sites. Six key genes--SRXN1, FURIN, IL18RAP, FOXO3, GCNT4, and FKBP5--were either up-regulated or down-regulated near hypomethylated CpG sites. These genes are associated with viral infiltration, immune activation, lung damage, and oxidative stress-related multi-organ failure, which are the hallmarks of severe COVID-19. Interestingly, during the recovery phase, methylation and gene expression levels returned to baseline, underscoring the rapid and reversible nature of these molecular changes. These findings provide insight into the dynamics of epigenetic and transcriptomic shifts according to the infectious stage, supporting potential prognostic and therapeutic approaches for severe COVID-19.

bioinformatics↗

Identification of 17 novel epigenetic biomarkers associated with anxiety disorders using differential methylation analysis followed by machine learning-based validation

BackgroundThe changes in DNA methylation patterns may reflect both physical and mental well-being, the latter being a relatively unexplored avenue in terms of clinical utility for psychiatric disorders. In this study, our objective was to identify the methylation-based biomarkers for anxiety disorders and subsequently validate their reliability. MethodsA comparative differential methylation analysis was performed on whole blood samples from 94 anxiety disorder patients and 296 control samples using targeted bisulfite sequencing. Subsequent validation of identified biomarkers employed an artificial intelligence- based risk prediction models: a linear calculation-based methylation risk score model and two tree-based machine learning models: Random Forest and XGBoost. Results17 novel epigenetic methylation biomarkers were identified to be associated with anxiety disorders. These biomarkers were predominantly localized near CpG islands, and they were associated with two distinct biological processes: 1) cell apoptosis and mitochondrial dysfunction and 2) the regulation of neurosignaling. We further developed a robust diagnostic risk prediction system to classify anxiety disorders from healthy controls using the 17 biomarkers. Machine learning validation confirmed the robustness of our biomarker set, with XGBoost as the best-performing algorithm, an area under the curve of 0.876. ConclusionOur findings support the potential of blood liquid biopsy in enhancing the clinical utility of anxiety disorder diagnostics. This unique set of epigenetic biomarkers holds the potential for early diagnosis, prediction of treatment efficacy, continuous monitoring, health screening, and the delivery of personalized therapeutic interventions for individuals affected by anxiety disorders.

bioinformatics↗

Korea4K: whole genome sequences of 4,157 Koreans with 107 phenotypes derived from extensive health check-ups

We present 4,157 whole-genome sequences (Korea4K) coupled with 107 health check-up parameters as the largest whole genomic resource of Koreans. Korea4K provides 45,537,252 variants and encompasses most of the common and rare variants in Koreans. We identified 1,356 new geno-phenotype associations which were not found by the previous Korea1K dataset. Phenomics analyses revealed 24 genetic correlations, 1,131 pleiotropic variants, and 127 causal relationships from Mendelian randomization. Moreover, the Korea4K imputation reference panel showed a superior imputation performance to Korea1K. Collectively, Korea4K provides the most extensive genomic and phenomic data resources for discovering clinically relevant novel genome-phenome associations in Koreans.

genomics↗