Search bioRxiv⌕ Search

Biology subjects

Tomizuka, K.

Publications and source records attributed to Tomizuka, K..

3 recordsLinked to original sources

Predicted constrained accessible regions mark regulatory elements and causal variants

Open chromatin regions (OCRs) define cell-type-specific regulatory elements across the genome, yet their functional significance varies, making it challenging to pinpoint biologically essential regions. Here, we introduce CAMBUS (Chromatin Accessibility Mutation Burden Score), a machine-learning framework that identifies active and evolutionarily constrained OCRs by leveraging surrounding DNA sequences. Applying CAMBUS to 29 immune cell types, we identified 66,043 constrained OCRs, which were substantially enriched in the known constraint genome (odds ratio=11.45 (95% confidence interval 9.33-14.05), P=4.7 x 10-68), while 90% of these OCRs were not prioritized by existing constraint metrics. These OCRs were highly enriched for known enhancers and super-enhancers, independent of known epigenetic markers and annotated regions, and overlapped with regulatory elements implicated in immune-mediated diseases and experimentally validated functional variants, including rare variants, particularly in leukocyte-related traits. Furthermore, CAMBUS revealed cell-type-specific transcriptional regulatory landscapes, linking genetic constraint with gene regulation in immune cells and identifying plausible connections between 1,533 causal variants and 70 complex traits. By defining biologically constrained regulatory elements at high resolution, CAMBUS provides a framework for understanding the selective pressures shaping the non-coding genome and its role in human health and disease.

genetics↗

Comparative analysis of trans-chromosomic rodent models reveals improved somatic hypermutation and class-switch recombination in rats

Humanized rodent models, especially humanization of genetic/genomic components involved in immunity have significantly advanced our understanding of human immune system. Here, we utilized trans-chromosomic (Tc) technology to generate a TC-mAb rat model that stably harbors a mouse artificial chromosome carrying full-length human immunoglobulin (Ig) heavy and kappa light chain genes (IGHK-NAC) in a rat Ig knockout background. In contrast with TC-mAb mice, serum human IgG concentration was found higher than IgM. Number of lymphocytes was recovered, and B cell population in the spleen was normal. Remarkably, repertoire analysis revealed similarities between the model and human PBMCs; somatic hypermutation and class-switch recombination also more closely resembled humans. Furthermore, immunization resulted in generation of antigen-specific human antibodies. Collectively, our strategy to generate both rat and mouse models through introduction of the identical IGHK-NAC offers unprecedented opportunities to comprehensively evaluate genomic regulation and its outcomes associated with genomic sequences and host-derived protein factors.

bioengineering↗

Characterizing the quality metric in genotype imputation

Large-scale imputation reference panels are now available and have contributed to efficient genome-wide association studies through genotype imputation. However, it is still under debate whether large-size multi-ancestry or small-size population-specific reference panels are the optimal choices for under-represented populations. We imputed genotypes of East Asian (EAS; 180k Japanese) subjects using the Trans-Omics for Precision Medicine (TOPMed) reference panel and found that the standard imputation quality metric (Rsq) substantially overestimated the dosage r2 (squared correlation between imputed dosage and true genotype). Variance component analysis of Rsq revealed that the increased imputed-genotype certainty (dosages closer to 0, 1, or 2) caused upward bias, indicating some systemic bias in the imputation. Through systematic simulations using different template switching rates ({theta} value) in the hidden Markov model, we uncovered that the lower {theta} value increased the imputed-genotype certainty and Rsq; however, dosage r2 was insensitive to the {theta} value, thereby causing a deviation. In simulated reference panels with different sizes and ancestral diversities, the {theta} value estimates from Minimac decreased with the size of a single ancestry and increased with the ancestral diversity. Thus, Rsq could overestimate or underestimate dosage r2 for a subpopulation in the multi-ancestry panel and the deviation represents different imputed-dosage distributions. Finally, despite the impact of {theta} value, distant ancestries in the reference panel contributed only a few additional variants passing a predefined Rsq threshold. We conclude that the {theta} value has a substantial impact on the imputed dosage and the imputation quality metric value.

bioinformatics↗