Search bioRxivSearch

Biology subjects

Raine, K. M.

Publications and source records attributed to Raine, K. M..

3 recordsLinked to original sources

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics

Framework For Quality Assessment Of Whole Genome, Cancer Sequences

Working with cancer whole genomes sequenced over a period of many years in different sequencing centres requires a validated framework to compare the quality of these sequences. The Pan-Cancer Analysis of Whole Genomes (PCAWG) of the International Cancer Genome Consortium (ICGC), a project a cohort of over 2800 donors provided us with the challenge of assessing the quality of the genome sequences. A non-redundant set of five quality control (QC) measurements were assembled and used to establish a star rating system. These QC measures reflect known differences in sequencing protocol and provide a guide to downstream analyses of these whole genome sequences. The resulting QC measures also allowed for exclusion samples of poor quality, providing researchers within PCAWG, and when the data is released for other researchers, a good idea of the sequencing quality. For a researcher wishing to apply the QC measures for their data we provide a Docker Container of the software used to calculate them. We believe that this is an effective framework of quality measures for whole genome, cancer sequences, which will be a useful addition to analytical pipelines, as it has to the PCAWG project.

genomics

Universal Patterns Of Selection In Cancer And Somatic Tissues

Cancer develops as a result of somatic mutation and clonal selection, but quantitative measures of selection in cancer evolution are lacking. We applied methods from evolutionary genomics to 7,664 human cancers across 29 tumor types. Unlike species evolution, positive selection outweighs negative selection during cancer development. On average, <1 coding base substitution/tumor is lost through negative selection, with purifying selection only detected for truncating mutations in essential genes in haploid regions. This allows exome-wide enumeration of all driver mutations, including outside known cancer genes. On average, tumors carry [~]4 coding substitutions under positive selection, ranging from <1/tumor in thyroid and testicular cancers to >10/tumor in endometrial and colorectal cancers. Half of driver substitutions occur in yet-to-be-discovered cancer genes. With increasing mutation burden, numbers of driver mutations increase, but not linearly. We identify novel cancer genes and show that genes vary extensively in what proportion of mutations are drivers versus passengers.\n\nHIGHLIGHTSO_LIUnlike the germline, somatic cells evolve predominantly by positive selection\nC_LIO_LINearly all ([~]99%) coding mutations are tolerated and escape negative selection\nC_LIO_LIFirst exome-wide estimates of the total number of driver coding mutations per tumor\nC_LIO_LI1-10 coding driver mutations per tumor; half occurring outside known cancer genes\nC_LI

genomics