Search bioRxivSearch

Biology subjects

Mungall, A. J.

Publications and source records attributed to Mungall, A. J..

6 recordsLinked to original sources

Comprehensive characterization of genomic, transcriptomic and epigenomic artifacts introduced in formalin-fixed, paraffin-embedded tissues.

Genomic, transcriptomic and epigenomic characterization has accelerated the discovery of clinically-relevant alterations in cancer, predominantly using fresh frozen (FF) specimens. However, clinical molecular pathology laboratories prefer formalin-fixed paraffin-embedded (FFPE) methods, known to introduce artifacts at the nucleic acid level, over fresh frozen methods. Extending the multi-platform analysis to FFPE specimens for comprehensive clinical molecular diagnosis requires a thorough understanding of the consequence of formalin-fixation. We present a detailed multi-platform characterization of FFPE preservation using paired FF specimens as the 'gold standard'. DNA and RNA were obtained from 38 patients across 6 cancer types using a FFPE optimized co-isolation. The impact of FFPE on exome sequencing was dependent on filtering, where a minimum coverage or supporting read filter can mitigate FFPE-specific false positives. Copy number alterations, MSI assessment, mutational signatures, and DNA methylation were comparable between FFPE and FF. FFPE biases in RNA expression can be overcome when using biology-relevant genes and we describe a novel consequence of FFPE on miRNA species diversity. Collectively, this data provides a broad view of FFPE artifact and offers best practices for overcome these biases.

bioinformatics

High-resolution structural genomics reveals new therapeutic vulnerabilities in glioblastoma

We investigated the role of 3D genome architecture in instructing functional properties of glioblastoma stem cells (GSCs) by generating the highest-resolution 3D genome maps to-date for this cancer. Integration of DNA contact maps with chromatin and transcriptional profiles identified specific mechanisms of gene regulation, including individual physical interactions between regulatory regions and their target genes. Residing in structurally conserved regions in GSCs was CD276, a gene known to play a role in immuno-modulation. We show that, unexpectedly, CD276 is part of a stemness network in GSCs and can be targeted with an antibody-drug conjugate to curb self-renewal, a key stemness property. Our results demonstrate that integrated structural genomics datasets can be employed to rationally identify therapeutic vulnerabilities in self-renewing cells.\n\nSIGNIFICANCEIn adult GBM, GSCs act as therapy-resistant reservoirs to nucleate tumor recurrence. New therapeutic approaches that target these cell populations hold the potential of significantly improving patient care and overall prognosis for this always-lethal cancer. Our work describes new links between 3D genome architecture and stemness properties in GSCs. In particular, through integration of multiple genomics and structural genomics datasets, we found an unexpected connection between immune-related genes and self-renewal programs in GBM. Among these, we show that targeting CD276 with knockdown strategies or specific antibody-drug conjugates achieve suppression of self-renewal. Strategies to target CD276+ cells are currently in clinical trials for solid tumors. Our results indicate that CD276-targeting agents could be deployed in GBM to specifically target GSC populations.\n\nHIGHLIGHTSO_LIWe generated high (sub-5 kb) resolution Hi-C maps for stem-like cells from GBM patients.\nC_LIO_LIIntegration of Hi-C and genomics datasets dissects mechanisms of gene regulation.\nC_LIO_LI3D genomes poise immune-related genes, including CD276, for expression.\nC_LIO_LITargeting CD276 curbs self-renewal properties of GBM cells.\nC_LI

cancer biology

Subclonal architecture, evolutionary trajectories and patterns of inheritance of germline variants in pediatric glioblastoma

Pediatric glioblastoma (pGBM) is a lethal cancer with no effective therapies. Intratumoral genetic heterogeneity and mode of tumor evolution have not been systematically addressed for this cancer. Whole-genome sequencing of germline-tumor pairs showed that pGBM is characterized by intratumoral genetic heterogeneity and consequent subclonal architecture. We found that pGBM undergoes extreme evolutionary trajectories, with primary and recurrent tumors having different subclonal compositions. Analysis of variant allele frequencies supported a model of tumor growth involving slow-cycling cancer stem cells that give rise to fast-proliferating progenitor-like cells and to non-dividing cells. pGBM patients germlines had subclonal structural variants, some of which underwent dynamic frequency fluctuations during tumor evolution. By sequencing germlines of mother-father-patient trios, we found that inheritance of deleterious germline variants from healthy parents cooperate with de novo germline and somatic events to the tumorigenic process. Our studies therefore challenge the current notion that pGBM is a relatively homogeneous molecular entity.

cancer biology

Resource: Scalable whole genome sequencing of 40,000 single cells identifies stochastic aneuploidies, genome replication states and clonal repertoires

Essential features of cancer tissue cellular heterogeneity such as negatively selected genome topologies, sub-clonal mutation patterns and genome replication states can only effectively be studied by sequencing single-cell genomes at scale and high fidelity. Using an amplification-free single-cell genome sequencing approach implemented on commodity hardware (DLP+) coupled with a cloud-based computational platform, we define a resource of 40,000 single-cell genomes characterized by their genome states, across a wide range of tissue types and conditions. We show that shallow sequencing across thousands of genomes permits reconstruction of clonal genomes to single nucleotide resolution through aggregation analysis of cells sharing higher order genome structure. From large-scale population analysis over thousands of cells, we identify rare cells exhibiting mitotic mis-segregation of whole chromosomes. We observe that tissue derived scWGS libraries exhibit lower rates of whole chromosome anueploidy than cell lines, and loss of p53 results in a shift in event type, but not overall prevalence in breast epithelium. Finally, we demonstrate that the replication states of genomes can be identified, allowing the number and proportion of replicating cells, as well as the chromosomal pattern of replication to be unambiguously identified in single-cell genome sequencing experiments. The combined annotated resource and approach provide a re-implementable large scale platform for studying lineages and tissue heterogeneity.

genomics

Genome-wide discovery of somatic coding and regulatory variants in Diffuse Large B-cell Lymphoma

Diffuse large B-cell lymphoma (DLBCL) is an aggressive cancer originating from mature B-cells. Many known driver mutations are over-represented in one of its two molecular subgroups, knowledge of which has aided in the development of therapeutics that target these features. The heterogeneity of DLBCL determined through prior genomic analysis suggests an incomplete understanding of its molecular aetiology, with a limited diversity of genetic events having thus far been attributed to the activated B-cell (ABC) subgroup. Through an integrative genomic analysis we uncovered genes and non-coding loci that are commonly mutated in DLBCL including putative regulatory sequences. We implicate recurrent mutations in the 3UTR of NFKBIZ as a novel mechanism of oncogene deregulation and found small amplifications associated with over-expression of FC-{gamma} receptor genes. These results inform on mechanisms of NF-{kappa}B pathway activation in ABC DLBCL and may reveal a high-risk population of patients that might not benefit from standard therapeutics.

genomics

It’s okay to be green: Draft genome of the North American Bullfrog (Rana [Lithobates] catesbeiana)

Frogs play important ecological roles as sentinels, insect control and food sources. Several species are important model organisms for scientific research to study embryogenesis, development, immune function, and endocrine signaling. The globally-distributed Ranidae (true frogs) are the largest frog family, and have substantial evolutionary distance from the model laboratory Xenopus frog species. Consequently, the extensive Xenopus genomic resources are of limited utility for Ranids and related frog species. More widely applicable amphibian genomic data is urgently needed as more than two-thirds of known species are currently threatened or are undergoing population declines.\n\nHerein, we report on the first genome sequence of a Ranid species, an adult male North American bullfrog (Rana [Lithobates] catesbeiana). We assembled high-depth Illumina reads (66-fold coverage), into a 5.8 Gbp (NG50 = 57.7 kbp) draft genome using ABySS v1.9.0. The assembly was scaffolded with LINKS and RAILS using pseudo-long-reads from targeted denovo assembler Kollector and Illumina Synthetic Long-Reads, as well as reads from long fragment (MPET) libraries. We predicted over 22,000 protein-coding genes using the MAKER2 pipeline and identified the genomic loci of 6,227 candidate long noncoding RNAs (IncRNAs) from a composite reference bullfrog transcriptome. Mitochondrial sequence analysis supported Lithobates as a subgenus of Rana. RNA-Seq experiments identified ~6,000 thyroid hormone- responsive transcripts in the back skin of premetamorphic tadpoles; the majority of which regulate DNA/RNA processing. Moreover, 1/6th of differentially-expressed transcripts were putative lncRNAs. Our draft bullfrog genome will serve as a useful resource for the amphibian research community.

genomics