Search bioRxiv⌕ Search

Biology subjects

Yang, J. O.

Publications and source records attributed to Yang, J. O..

4 recordsLinked to original sources

Telomere-to-Telomere Accurate and Gapless Korean Standard Reference Genome

We present KOREF1-G-TTAGGA, the first Telomere-to-Telomere Accurate and Gapless Genome Assembly, standing as the Korean standard reference genome. The paternal and maternal haplotypes spanned 2.91 and 3.03 Gb. Genome-wide, at least 99.15% of assembled sequences were reliably haplotype-resolved, and over 95% remained accurate even in the most error-prone loci, including centromeric satellite arrays and segmental duplications. Rare-k-mer copy-number concordance within satellite arrays held up 98.9% and 98.8% per haplotype. Bionano optical maps fully spanned all canonical rDNA arrays on the five acrocentric chromosomes. The assembly quality index (AQI) of 99.77 and 99.69 exceeded the reference-quality threshold of 90, and more than 99.99% of gene and cCRE sequence was free of structural error. Both haplotypes further showed high base-level accuracy, with consensus quality value (QV) of 81.19 and 79.03, corresponding to one error per 131 and 80 Mb. KOREF1-G-TTAGGA is among the highest-quality East Asian telomere-to-telomere assemblies. It moreover anchors a decade-spanning multi-ome reference dataset for defining individual molecular states. Together, these resources define the personal referenceome as a foundation for individual biology and precision medicine, with the assemblies, annotations, and all multi-omic data openly available at https://koreanreference.org.

genomics↗

10,239 whole genomes with multiomic and clinical health information as the Korean population multiomic reference dataset

We present Korea10K, the largest genomic dataset of the Korean population, comprising 10,239 high-coverage whole genomes (mean depth 30x) with matched multiomic profiles and phenotype data. Korea10K achieves complete and near-complete discovery of very rare and ultra-rare alleles, respectively, at 9,000 Korean genomes. This dataset provides the high-quality population-specific imputation panel, enabling accurate inference of low-frequency variants. Admixture analyses confirm the genetic homogeneity of the Korean population, despite its diverse Y-chromosomal, mitochondrial, and HLA repertoires. This pattern reflects a long and continuous lineage history characterized by persistent internal admixture and genomic homogenization over thousands of years on the Korean peninsula. We also identified 16.8 million genomic variants that directly modify CG sites by creating or abolishing CG dinucleotides, providing the population-scale evidence of coordinated genomic-epigenomic regulatory mechanism in Koreans.

genomics↗

KoNA: Korean Nucleotide Archive as a New Data Repository for Nucleotide Sequence Data

During the last decade, generation and accumulation of petabase-scale high-throughput sequencing data have resulted in ethical and technical challenges, including access to human data, and transfer, storage, and sharing of enormous amount of data. To promote data-driven research in biology, the Korean government announced that all the biological data generated from government-funded research projects should be deposited in the Korea BioData Station (K-BDS), which consists of multiple databases for individual data types. We introduce the Korean Nucleotide Archive (KoNA), a repository for nucleotide sequence data. As of July 2022, the Korean Read Archive in KoNA has collected over 477 TB of raw next generation sequencing data from several national genome projects. To ensure data quality and prepare for international alignment, a standard operating procedure (SOP) was adopted, which is similar to the International Nucleotide Sequence Database Collaboration. The SOP includes quality control processes for submitted data and metadata using an automated pipeline followed by manual examination. To ensure fast and stable data transfer, a high-speed transmission system called GBox is used in KoNA. Furthermore, the data uploaded to or downloaded from KoNA through GBox can be readily processed in a cloud-computing service for genomic data analysis called Bio-Express. This seamless coupling of KoNA, GBox, and Bio-Express enhances data experience including submission, access, and analysis of raw nucleotide sequences. KoNA not only satisfies the unmet needs for a national sequence repository in Korea, but also provides datasets to researchers globally and contribute to advances in genomics. KoNA is available at https://www.kobic.re.kr/kona/.

genomics↗

Metastatic prognostic ability of lung cancer stromal cells from a single-cell RNA-seq perspective

Single-cell RNA sequencing (scRNA-seq) has been widely studied and analyzed to understand cancer heterogeneity. Metastasis and invasion through communication with immune cells have been widely studied; stromal cells are known to change during cancer progression and cause metastasis, but little is known about their inherent metastasis prognostic abilities. This study investigated the abilities of stromal cells by analyzing the scRNA-seq data of adjacent, tumor, and metastasized tissues, biopsied from 15 patients. We considered fibroblast and smooth muscle cells as cell subtypes of stromal cells. We detected decorin (DCN) and insulin-like growth factor-binding protein 7 (IGFBP7) as sub-cell type markers conserved in tumor and metastasis. We found an organic relationship that affects metastasis by assessing the interaction between the expression and related pathways among the assigned stromal cell subtypes. In addition to the interactions of sub-cell-types within stromal cells, we also studied communication between stromal cells and the five assigned lung cancer cell types, and observed its relation with migration and metastasis; the role of DCN as a mediating factor was also studied. Results of our study indicated that DCN and IGFBP7 are factors that can be monitored in the follow-up of prognostic metastasis factors in patients with lung cancer. Therefore, DCN and IGFBP7, which are the assigned sub-cell types marker in lung cancer stromal cells, can be used as potential biomarkers for follow-up in lung cancer metastasis. Our study assigned stromal cell subtypes of lung cancer and detected markers that can be used as monitoring factors during metastasis of lung cancer. This suggests that DCN and IGFBP7 are potential biomarkers to evaluate metastatic ability and follow-up as metastatic prognostic factors in lung cancer patients. Author SummaryLung cancer is known to be a disease with high heterogeneity. In particular, it has been studied that stromal cells of lung cancer are involved in metastasis in cancer. Finding specific markers of stromal cell subtypes can lead to more accurate cancer markers. In this study, we assigned stromal cell subtypes as FB and SMC through scRNA-seq data analysis. DCN and IGFBP7, which are markers specifically expressed in stromal cell sub-cell types of NSCLC and metastatic NSCLC, were found. Whether DCN and IGFBP7 can be used as prognostic factors for metastasis was verified through co-expression analysis of FB and SMC and cell-cell communication. Specific expression as a marker was verified by confirming the role in the hub-gene network and the target genes of communication. Our findings suggest that DCN and IGFBP7 are potential biomarkers for evaluating the metastasis prognostic ability of stromal cells and can be monitored in the follow-up of metastasis prognostic factors in lung cancer patients.

bioinformatics↗