Search bioRxivSearch

Biology subjects

Campbell, P.

Publications and source records attributed to Campbell, P..

8 recordsLinked to original sources

Process-specific somatic mutation distributions vary with three-dimensional genome structure

Somatic mutations arise during the life history of a cell. Mutations occurring in cancer driver genes may ultimately lead to the development of clinically detectable disease. Nascent cancer lineages continue to acquire somatic mutations throughout the neoplastic process and during cancer evolution (Martincorena and Campbell, 2015). Extrinsic and endogenous mutagenic factors contribute to the accumulation of these somatic mutations (Zhang and Pellman, 2015). Understanding the underlying factors generating somatic mutations is crucial for developing potential preventive, therapeutic and clinical decisions. Earlier studies have revealed that DNA replication timing (Stamatoyannopoulos et al., 2009) and chromatin modifications (Schuster-Bockler and Lehner, 2012) are associated with variations in mutational density. What is unclear from these early studies, however, is whether all extrinsic and exogenous factors that drive somatic mutational processes share a similar relationship with chromatin state and structure. In order to understand the interplay between spatial genome organization and specific individual mutational processes, we report here a study of 3000 tumor-normal pair whole genome datasets from more than 40 different human cancer types. Our analyses revealed that different mutational processes lead to distinct somatic mutation distributions between chromatin folding domains. APOBEC- or MSI-related mutations are enriched in transcriptionally-active domains while mutations occurring due to tobacco-smoke, ultraviolet (UV) light exposure or a signature of unknown aetiology (signature 17) enrich predominantly in transcriptionally-inactive domains. Active mutational processes dictate the mutation distributions in cancer genomes, and we show that mutational distributions shift during cancer evolution upon mutational processes switch. Moreover, a dramatic instance of extreme chromatin structure in humans, that of the unique folding pattern of the inactive X-chromosome leads to distinct somatic mutation distribution on X chromosome in females compared to males in various cancer types. Overall, the interplay between three-dimensional genome organization and active mutational processes has a substantial influence on the large-scale mutation rate variations observed in human cancer.

genomics

SNP Variable Selection by Generalized Graph Domination

High-throughput sequencing technology has revolutionized both medical and biological research by generating exceedingly large numbers of genetic variants. The resulting datasets share a number of common characteristics that might lead to poor generalization capacity. Concerns include noise accumulated due to the large number of predictors, sparse information regarding the p >> n problem, and overfitting and model mis-identification resulting from spurious collinearity. Additionally, complex correlation patterns are present among variables. As a consequence, reliable variable selection techniques play a pivotal role in predictive analysis, generalization capability, and robustness in clustering, as well as interpretability of the derived models.\n\nK-dominating set, a parameterized graph-theoretic generalization model, was used to model SNP (single nucleotide polymorphism) data as a similarity network and searched for representative SNP variables. In particular, each SNP was represented as a vertex in the graph, (dis)similarity measures such as correlation coefficients or pairwise linkage disequilibrium were estimated to describe the relationship between each pair of SNPs; a pair of vertices are adjacent, i.e. joined by an edge, if the pairwise similarity measure exceeds a user-specified threshold. A minimum K-dominating set in the SNP graph was then made as the smallest subset such that every SNP that is excluded from the subset has at least k neighbors in the selected ones. The strength of\n\nk-dominating set selection in identifying independent variables, and in culling representative variables that are highly correlated with others, was demonstrated by a simulated dataset. The advantages of k-dominating set variable selection were also illustrated in two applications: pedigree reconstruction using SNP profiles of 1,372 Douglas-fir trees, and species delineation for 226 grasshopper mouse samples. A C++ source code that implements SNP-SELECT and uses Gurobi optimization solver for the k-dominating set variable selection is available (https://github.com/transgenomicsosu/SNP-SELECT).

genetics

Back pain, mental health and substance use are associated in adolescents

BackgroundDuring adolescence, prevalence of pain and health risk factors such as smoking, alcohol use, and poor mental health rise sharply. While these risk factors and mental health are accepted public health concerns, the same is not true for pain. The aim of this study was to describe the relationship between back pain and health risk factors in adolescents.\n\nMethodsCross-sectional data from the Healthy Schools Healthy Futures study, and the Australian Child Wellbeing Project was used. The mean age of participants was 14-15 years. Children were stratified according to the frequency they experienced back pain over the past 6 months. Within each strata, the proportion of children that reported drinking alcohol or smoking in the past month and the proportion that experienced feelings of anxiety or depression was reported. Test-for-trend analyses assessed whether increasing frequency of pain was associated with health risk factors.\n\nResultsData from approximately 2,500 and 3,900 children in the two studies was analysed. Larger proportions of children smoked or drank alcohol within each strata of increasing pain frequency. The trend with report of anxiety and depression was less clear, although there was a marked difference between the children that reported pain rarely or never, and those that experienced back pain more frequently.\n\nConclusionTwo large, independent samples show Australian adolescents that experience back pain more frequently are also more likely to smoke, drink alcohol and report feelings of anxiety and depression. Pain appears to be part of the picture of general health risk in adolescents.\n\nWhat is already known on this subject?The prevalence of back pain rises steeply during the adolescent years, and is responsible for considerable personal impact in a substantial minority. During this time, indicators of adverse health risk such as smoking, alcohol use, anxiety and depression also increase in prevalence. Pain and lifestyle-related health risk factors can have ongoing consequences that stretch into adulthood.\n\nWhat this study adds?This study shows a close relationship between increasing pain frequency, and tendency to engage in health risk behaviours and experience indicators of poor mental health in adolescents. This study shows that pain may be an important consideration in understanding the general health, and health risk in adolescents.

epidemiology

Selective and mechanistic sources of recurrent rearrangements across the cancer genome

Cancer cells can acquire profound alterations to the structure of their genomes, including rearrangements that fuse distant DNA breakpoints. We analyze the distribution of somatic rearrangements across the cancer genome, using whole-genome sequencing data from 2,693 tumor-normal pairs. We observe substantial variation in the density of rearrangement breakpoints, with enrichment in open chromatin and sites with high densities of repetitive elements. After accounting for these patterns, we identify significantly recurrent breakpoints (SRBs) at 52 loci, including novel SRBs near BRD4 and AKR1C3. Taking into account both loci fused by a rearrangement, we observe different signatures resembling either single breaks followed by strand invasion or two separate breaks that become joined. Accounting for these signatures, we identify 90 pairs of loci that are significantly recurrently juxtaposed (SRJs). SRJs are primarily tumor-type specific and tend to involve genes with tissue-specific expression. SRJs were frequently associated with disruption of topology-associated domains, juxtaposition of enhancer elements, and increased expression of neighboring genes. Lastly, we find that the power to detect SRJs decreases for short rearrangements, and that reliable detection of all driver SRJs will require whole-genome sequencing data from an order of magnitude more cancer samples than currently available.

cancer biology

Patterns of structural variation in human cancer

A key mutational process in cancer is structural variation, in which rearrangements delete, amplify or reorder genomic segments ranging in size from kilobases to whole chromosomes. We developed methods to group, classify and describe structural variants, applied to >2,500 cancer genomes. Nine signatures of structural variation emerged. Deletions have trimodal size distribution; assort unevenly across tumour types and patients; enrich in late-replicating regions; and correlate with inversions. Tandem duplications also have trimodal size distribution, but enrich in early-replicating regions, as do unbalanced translocations. Replication-based mechanisms of rearrangement generate varied chromosomal structures with low-level copy number gains and frequent inverted rearrangements. One prominent structure consists of 1-7 templates copied from distinct regions of the genome strung together within one locus. Such cycles of templated insertions correlate with tandem duplications, frequently activating the telomerase gene, TERT, in liver cancer. Cancers access many rearrangement processes, flexibly sculpting the genome to maximise oncogenic potential.

cancer biology

Framework For Quality Assessment Of Whole Genome, Cancer Sequences

Working with cancer whole genomes sequenced over a period of many years in different sequencing centres requires a validated framework to compare the quality of these sequences. The Pan-Cancer Analysis of Whole Genomes (PCAWG) of the International Cancer Genome Consortium (ICGC), a project a cohort of over 2800 donors provided us with the challenge of assessing the quality of the genome sequences. A non-redundant set of five quality control (QC) measurements were assembled and used to establish a star rating system. These QC measures reflect known differences in sequencing protocol and provide a guide to downstream analyses of these whole genome sequences. The resulting QC measures also allowed for exclusion samples of poor quality, providing researchers within PCAWG, and when the data is released for other researchers, a good idea of the sequencing quality. For a researcher wishing to apply the QC measures for their data we provide a Docker Container of the software used to calculate them. We believe that this is an effective framework of quality measures for whole genome, cancer sequences, which will be a useful addition to analytical pipelines, as it has to the PCAWG project.

genomics

The Human Cell Atlas

The recent advent of methods for high-throughput single-cell molecular profiling has catalyzed a growing sense in the scientific community that the time is ripe to complete the 150-year-old effort to identify all cell types in the human body, by undertaking a Human Cell Atlas Project as an international collaborative effort. The aim would be to define all human cell types in terms of distinctive molecular profiles (e.g., gene expression) and connect this information with classical cellular descriptions (e.g., location and morphology). A comprehensive reference map of the molecular state of cells in healthy human tissues would propel the systematic study of physiological states, developmental trajectories, regulatory circuitry and interactions of cells, as well as provide a framework for understanding cellular dysregulation in human disease. Here we describe the idea, its potential utility, early proofs-of-concept, and some design considerations for the Human Cell Atlas.

cell biology

Genome-wide detection of structural variants and indels by local assembly

Structural variants (SVs), including small insertion and deletion variants (indels), are challenging to detect through standard alignment-based variant calling methods. Sequence assembly offers a powerful approach to identifying SVs, but is difficult to apply at-scale genome-wide for SV detection due to its computational complexity and the difficulty of extracting SVs from assembly contigs. We describe SvABA, an efficient and accurate method for detecting SVs from short-read sequencing data using genome-wide local assembly with low memory and computing requirements. We evaluated SvABAs performance on the NA12878 human genome and in simulated and real cancer genomes. SvABA demonstrates superior sensitivity and specificity across a large spectrum of SVs, and substantially improved detection performance for variants in the 20-300 bp range, compared with existing methods. SvABA also identifies complex somatic rearrangements with chains of short (< 1,000 bp) templated-sequence insertions copied from distant genomic regions. We applied SvABA to 344 cancer genomes from 11 cancer types, and found that templated-sequence insertions occur in ~4% of all somatic rearrangements. Finally, we demonstrate that SvABA can identify sites of viral integration and cancer driver alterations containing medium-sized SVs.

genomics