Search bioRxivSearch

Biology subjects

Nakagawa, H.

Publications and source records attributed to Nakagawa, H..

9 recordsLinked to original sources

Comprehensive Analysis of Indels in Whole-genome Microsatellite Regions and Microsatellite Instability across 21 Cancer Types

Microsatellites are repeats of 1-6bp units and [~]10 million microsatellites have been identified across the human genome. Microsatellites are vulnerable to DNA mismatch errors, and have thus been used to detect cancers with mismatch repair deficiency. To reveal the mutational landscape of the microsatellite repeat regions at the genome level, we analyzed approximately 20.1 billion microsatellites in 2,717 whole genomes of pan-cancer samples across 21 tissue types. Firstly, we developed a new insertion and deletion caller (MIMcall) that takes into consideration the error patterns of different types of microsatellites. Among the 2,717 pan-cancer samples, our analysis identified 31 samples, including colorectal, uterus, and stomach cancers, with higher microsatellite mutation rate ([≥] 0.03), which we defined as microsatellite instability (MSI) cancers in genome-wide level. Next, we found 20 highly-mutated microsatellites that can be used to detect MSI cancers with high sensitivity. Third, we found that replication timing and DNA shape were significantly associated with mutation rates of the microsatellites. Analysis of germline variation of the microsatellites suggested that the amount of germline variations and somatic mutation rates were correlated. Lastly, analysis of mutations in mismatch repair genes showed that somatic SNVs and short indels had larger functional impact than germline mutations and structural variations. Our analysis provides a comprehensive picture of mutations in the microsatellite regions, and reveals possible causes of mutations, as well as provides a useful marker set for MSI detection.

cancer biology

ALPHLARD: a Bayesian method for analyzing HLA genes from whole genome sequence data

Although human leukocyte antigen (HLA) genotyping based on amplicon, whole exome sequence (WES), and RNA sequence data has been achieved in recent years, accurate genotyping from whole genome sequence (WGS) data remains a challenge due to the low depth. Furthermore, there is no method to identify the sequences of unknown HLA types not registered in HLA databases. We developed a Bayesian model, called ALPHLARD, that collects reads potentially generated from HLA genes and accurately determines a pair of HLA types for each of HLA-A, -B, -C, -DPA1, -DPB1, -DQA1, -DQB1, and -DRB1 genes at 6-digit resolution. Furthermore, ALPHLARD can detect rare germline variants not stored in HLA databases and call somatic mutations from paired normal and tumor sequence data. We illustrate the capability of ALPHLARD using 253 WES data and 25 WGS data from Illumina platforms. By comparing the results of HLA genotyping from SBT and amplicon sequencing methods, ALPHLARD achieved 98.8% for WES data and 98.5% for WGS data at 4-digit resolution. We also detected three somatic point mutations and one case of loss of heterozygosity in the HLA genes from the WGS data. ALPHLARD showed good performance for HLA genotyping even from low-coverage data. It also has a potential to detect rare germline variants and somatic mutations in HLA genes. It would help to fill in the current gaps in HLA reference databases and unveil the immunological significance of somatic mutations identified in HLA genes.

bioinformatics

Immuno-genomic PanCancer Landscape Reveals Diverse Immune Escape Mechanisms and Immuno-Editing Histories

Immune reactions in the tumor micro-environment are one of the cancer hallmarks and emerging immune therapies have been proven effective in many types of cancer. To investigate cancer genome-immune interactions and the role of immuno-editing or immune escape mechanisms in cancer development, we analyzed 2,834 whole genomes and RNA-seq datasets across 31 distinct tumor types from the PanCancer Analysis of Whole Genomes (PCAWG) project with respect to key immuno-genomic aspects. We show that selective copy number changes in immune-related genes could contribute to immune escape. Furthermore, we developed an index of the immuno-editing history of each tumor sample based on the information of mutations in exonic regions and pseudogenes. Our immuno-genomic analyses of pan-cancer analyses have the potential to identify a subset of tumors with immunogenicity and diverse background or intrinsic pathways associated with their immune status and immuno-editing history.

genomics

Genetic structure and sex-biased gene flow in the history of southern African populations

ObjectivesWe investigated the genetic history of southern African populations with a special focus on their paternal history. We reexamined previous claims that the Y-chromosome haplogroup E1b1b was brought to southern Africa by pastoralists from eastern Africa, and investigated patterns of sex-biased gene flow in southern Africa.\n\nMaterial and MethodsWe analyzed previously published complete mtDNA genome sequences and ~900 kb of NRY sequences from 23 populations from Namibia, Botswana and Zambia, as well as haplogroup frequencies from a large sample of southern African populations and 23 newly genotyped Y-linked STR loci for samples assigned to haplogroup E1b1b.\n\nResultsOur results support an eastern African origin for Y-chromosome haplogroup E1b1b; however, its current distribution in southern Africa is not strongly associated with pastoralism, suggesting a more complex origin for pastoralism in this region. We confirm that the Bantu expansion had a notable genetic impact in southern Africa, and that in this region it was probably a rapid, male-dominated expansion. Furthermore, we find a significant increase in the intensity of sex-biased gene flow from north to south, which may reflect changes in the social dynamics between Khoisan and Bantu groups over time.\n\nConclusionsOur study shows that the population history of southern Africa has been very complex, with different immigrating groups mixing to different degrees with the autochthonous populations. The Bantu expansion led to heavily sex-biased admixture as a result of interactions between Khoisan females and Bantu males, with a geographic gradient which may reflect changes in the social dynamics between Khoisan and Bantu groups over time.

genetics

Quantifying Immune-Based Counterselection of Somatic Mutations

It is now well established that somatic mutations in protein-coding regions can generate neoantigens, and that these can be recognized by the immune system and contribute to clearance of developing cancers. However, there is currently no model that can quantitatively predict the neoantigenic effect of any given somatic mutation. Here, we examined signatures of immune selection pressure on the distribution of somatic mutations. We quantified the extent to which somatic mutations are significantly depleted in peptides that are predicted to be displayed by major histocompatibility complex (MHC) class I proteins. We characterized the dependence of this depletion on expression level. We then examined whether immune selection pressure on somatic mutations changes depending on whether the patient had either one or two MHC-encoding alleles that can display the peptide. Our results indicate that MHC-encoding alleles are, in general, incompletely dominant, i.e., that having two copies of the display-enabling allele is more effective in displaying that peptide than having just one copy. More generally, a quantitative understanding of counter-selection of identifiable subclasses of neoantigenic somatic variation could guide immunotherapy or aid in developing personalized cancer vaccines.

genetics

Germline determinants of the somatic mutation landscape in 2,642 cancer genomes

Cancers develop through somatic mutagenesis, however germline genetic variation can markedly contribute to tumorigenesis via diverse mechanisms. We discovered and phased 88 million germline single nucleotide variants, short insertions/deletions, and large structural variants in whole genomes from 2,642 cancer patients, and employed this genomic resource to study genetic determinants of somatic mutagenesis across 39 cancer types. Our analyses implicate damaging germline variants in a variety of cancer predisposition and DNA damage response genes with specific somatic mutation patterns. Mutations in the MBD4 DNA glycosylase gene showed association with elevated C>T mutagenesis at CpG dinucleotides, a ubiquitous mutational process acting across tissues. Analysis of somatic structural variation exposed complex rearrangement patterns, involving cycles of templated insertions and tandem duplications, in BRCA1-deficient tumours. Genome-wide association analysis implicated common genetic variation at the APOBEC3 gene cluster with reduced basal levels of somatic mutagenesis attributable to APOBEC cytidine deaminases across cancer types. We further inferred over a hundred polymorphic L1/LINE elements with somatic retrotransposition activity in cancer. Our study highlights the major impact of rare and common germline variants on mutational landscapes in cancer.

genomics

In-depth characterization of the cisplatin mutational signature in a human cell line and in esophageal and liver tumors

Background and aimsCisplatin reacts with DNA, and thereby likely generates a characteristic pattern of somatic mutations, called a mutational signature. Despite widespread use of cisplatin in cancer treatment and its role in contributing to secondary malignancies, its mutational signature has not been delineated. We hypothesize that cisplatins mutational signature can serve as a biomarker to identify cisplatin mutagenesis in suspected secondary malignancies. Knowledge of which tissues are at risk of developing cisplatin-induced secondary malignancies could lead to guidelines for non-invasive monitoring for secondary malignancies after cisplatin chemotherapy.\n\nMethodsWe performed whole genome sequencing of 10 independent clones of cisplatin-exposed MCF-10A and HepG2 cells, and delineated the patterns of single- and dinucleotide mutations in terms of flanking sequence, transcription strand bias, and other characteristics. We used the mSigAct signature presence test and non-negative matrix factorization to search for cisplatin mutagenesis in hepatocellular carcinomas and esophageal adenocarcinomas.\n\nResultsAll clones showed highly consistent patterns of single- and dinucleotide substitutions. The proportion of dinucleotide substitutions was high: 8.1% of single nucleotide substitutions were part of dinucleotide substitutions, presumably due to cisplatins propensity to form intra-and inter-strand crosslinks between purine bases in DNA. We identified likely cisplatin exposure in 9 hepatocellular carcinomas and 3 esophageal adenocarcinomas. All hepatocellular carcinomas for which clinical data were available and all esophageal cancers indeed had histories of cisplatin treatment.\n\nConclusionsWe experimentally delineated the single- and dinucleotide mutational signature of cisplatin. This signature enabled us to detect previous cisplatin exposure in human hepatocellular carcinomas and esophageal adenocarcinomas with high confidence.

cancer biology

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics

Comprehensive Molecular Characterization of Mitochondrial Genomes in Human Cancers

Mitochondria are essential cellular organelles that play critical roles in cancer development. Through International Cancer Genome Consortium, we performed a multidimensional characterization of mitochondrial genomes using the whole-genome sequencing data of ~2,700 patients across 37 cancer types and related RNA-sequencing data. Our analysis presents the most definitive mutational landscape of mitochondrial genomes including a novel hypermutated case. We observe similar mutational signatures across cancer types, suggesting powerful endogenous mutational processes in mitochondria. Truncating mutations are remarkably enriched in kidney, colorectal and thyroid cancers and associated with the activation of critical signaling pathways. We find frequent somatic nuclear transfers of mitochondrial DNA (especially in skin and lung cancers), some of which disrupt therapeutic target genes (e.g., ERBB2). The mitochondrial DNA copy number shows great variations within and across cancers and correlates with clinical variables. Co-expression analysis highlights the function of mitochondrial genes in oxidative phosphorylation, DNA repair, and cell cycle; and reveals their connections with clinically actionable genes. Our study, including an open-access data portal, lays a foundation for understanding the interplays between the cancer mitochondrial and nuclear genomes and translating mitochondrial biology into clinical applications.

genomics