Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Creating Standards for Evaluating Tumour Subclonal Reconstruction

Tumours evolve through time and space. Computational techniques have been developed to infer their evolutionary dynamics from DNA sequencing data. A growing number of studies have used these approaches to link molecular cancer evolution to clinical progression and response to therapy. There has not yet been a systematic evaluation of methods for reconstructing tumour subclonality, in part due to the underlying mathematical and biological complexity and to difficulties in creating gold-standards. To fill this gap, we systematically elucidated the key algorithmic problems in subclonal reconstruction and developed mathematically valid quantitative metrics for evaluating them. We then created approaches to simulate realistic tumour genomes, harbouring all known mutation types and processes both clonally and subclonally. We then simulated 580 tumour genomes for reconstruction, varying tumour read-depth and benchmarking somatic variant detection and subclonal reconstruction strategies. The inference of tumour phylogenies is rapidly becoming standard practice in cancer genome analysis; this study creates a baseline for its evaluation.

bioinformatics

Comparing alternative pipelines for cross-platform microarray gene expression data integration with RNA-seq data in breast cancer

BackgroundAccording to major public repositories statistics an overwhelming majority of the existing and newly uploaded data originates from microarray experiments. Unfortunately, the potential of this data to bring new insights is limited by the effects of individual study-specific biases due to small number of biological samples. Increasing sample size by direct microarray data integration increases the statistical power to obtain a more precise estimate of gene expression in a population of individuals resulting in lower false discovery rates. However, despite numerous recommendations for gene expression data integration, there is a lack of a systematic comparison of different processing approaches aimed to asses microarray platforms diversity and ambiguous probesets to genes correspondence, leading to low number of studies applying integration.\n\nResultsHere, we investigated five different approaches of the microarrays data processing in comparison with RNA-seq data on breast cancer samples. We aimed to evaluate different probesets annotations as well as different procedures of choosing between probesets mapped to the same gene. We show that pipelines rankings are mostly preserved across Affymetrix and Illumina platforms. BrainArray approach based on updated annotation and redesigned probesets definition and choosing probeset with the maximum average signal across the samples have best correlation with RNA-seq, while averaging probesets signals as well as scoring the quality of probes sequences mapping to the transcripts of the targeted gene have worse correlation. Finally, randomly selecting probeset among probesets mapped to the same gene significantly decreases the correlation with RNA-seq.\n\nConclusionWe show that methods, which rely on actual probesets signal intensities, are advantageous to methods considering biological characteristics of the probes sequences only and that cross-platform integration of datasets improves correlation with the RNA-seq data. We consider the results obtained in this paper contributive to the integrative analysis as a worthwhile alternative to the classical meta-analysis of the multiple gene expression datasets.

Bioinformatics

Independent component analysis provides clinically relevant insights into the biology of melanoma patients

The integration of publicly available and new patient-derived transcriptomic datasets is not straightforward and requires specialized approaches to deal with heterogeneity at technical and biological levels. Here we present a methodology that can overcome technical biases, predict clinically relevant outcomes and identify tumour-related biological processes in patients using previously collected large reference datasets. The approach is based on independent component analysis (ICA) - an unsupervised method of signal deconvolution. We developed parallel consensus ICA that robustly decomposes merged new and reference datasets into signals with minimal mutual dependency. By applying the method to a small cohort of primary melanoma and control samples combined with a large public melanoma dataset, we demonstrate that our method distinguishes cell-type specific signals from technical biases and allows to predict clinically relevant patient characteristics. Cancer subtypes, patient survival and activity of key tumour-related processes such as immune response, angiogenesis and cell proliferation were characterized. Additionally, through integration of transcriptomes and miRNomes, the method identified biological functions of miRNAs, which would otherwise not be possible.

genomics

EMT network-based feature selection improves prognosis prediction in lung adenocarcinoma

Various feature selection algorithms have been proposed to identify cancer prognostic biomarkers. In recent years, however, their reproducibility is criticized. The performance of feature selection algorithms is shown to be affected by the datasets, underlying networks and evaluation metrics. One of the causes is the curse of dimensionality, which makes it hard to select the features that generalize well on independent data. Even the integration of biological networks does not mitigate this issue because the networks are large and many of their components are not relevant for the phenotype of interest. With the availability of multi-omics data, integrative approaches are being developed to build more robust predictive models. In this scenario, the higher data dimensions create greater challenges.\n\nWe proposed a phenotype relevant network-based feature selection (PRNFS) framework and demonstrated its advantages in lung cancer prognosis prediction. We constructed cancer prognosis relevant networks based on epithelial mesenchymal transition (EMT) and integrated them with different types of omics data for feature selection. With less than 2.5% of the total dimensionality, we obtained EMT prognostic signatures that achieved remarkable prediction performance (average AUC values >0.8), very significant sample stratifications, and meaningful biological interpretations. In addition to finding EMT signatures from different omics data levels, we combined these single-omics signatures into multi-omics signatures, which improved sample stratifications significantly. Both single- and multi-omics EMT signatures were tested on independent multi-omics lung cancer datasets and significant sample stratifications were obtained.

bioinformatics

Predicting Carriers of Ongoing Selective Sweeps Without Knowledge of the Favored Allele

Methods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying many selective sweeps. For most of these sweeps, the favored allele remains unknown, making it difficult to distinguish carriers of the sweep from non-carriers. In an ongoing selective sweep, carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones. The main contribution of this paper is the development and analysis of a new statistic, the Haplotype Allele Frequency (HAF) score. The HAF score, assigned to individual haplotypes in a sample, naturally captures many of the properties shared by haplotypes carrying a favored allele. We provide a theoretical framework for computing expected HAF scores under different evolutionary scenarios, and we validate the theoretical predictions with simulations. As an application of HAF score computations, we develop an algorithm (PreCIOSS: Predicting Carriers of Ongoing Selective Sweeps) to identify carriers of the favored allele in selective sweeps, and we demonstrate its power on simulations of both hard and soft sweeps, as well as on data from well-known sweeps in human populations.\n\nAuthor summaryMethods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying genomic regions under positive selection. However, methods that detect positive selective sweeps do not typically identify the favored allele, or even the haplotypes carrying the favored allele. The main contribution of this paper is the development and analysis of a new statistic (the HAF score), assigned to individual haplotypes. Using both theoretical analyses and simulations, we describe how the HAF scores differ for carriers and non-carriers of the favored allele, and how they change dynamically during a selective sweep. We also develop an algorithm, PreCIOSS, for separating carriers and non-carriers. Our tool has broad applicability as carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones.

Evolutionary Biology

CD44 Controls Endothelial Proliferation and Functions as Endogenous Inhibitor of Angiogenesis

CD44 transmembrane glycoprotein is involved in angiogenesis, but it is not clear whether CD44 functions as a pro- or antiangiogenic molecule. Here, we assess the role of CD44 in angiogenesis and endothelial proliferation by using Cd44-null mice and CD44 silencing in human endothelial cells. We demonstrate that angiogenesis is increased in Cd44-null mice compared to either wild-type or heterozygous animals. Silencing of CD44 expression in cultured endothelial cells results in their augmented proliferation and viability. The growth-suppressive effect of CD44 is mediated by its extracellular domain and is independent of its hyaluronan binding function. CD44-mediated effect on cell proliferation is independent of specific angiogenic growth factor stimulation. These results show that CD44 expression on endothelial cells constrains endothelial cell proliferation and angiogenesis. Thus, endothelial CD44 might serve as a therapeutic target both in the treatment of cardiovascular diseases, where endothelial protection is desired, as well as in cancer treatment, due to its antiangiogenic properties.

Cell Biology

QuickRNASeq: Guide For Pipeline Implementation And For Interactive Results Visualization

i.Summary/AbstractSequencing of transcribed RNA molecules (RNA-seq) has been used wildly for studying cell transcriptomes in bulk or at the single-cell level (1, 2, 3) and is becoming the de facto technology for investigating gene expression level changes in various biological conditions, on the time course, and under drug treatments. Furthermore, RNA-Seq data helped identify fusion genes that are related to certain cancers (4). Differential gene expression before and after drug treatments provides insights to mechanism of action, pharmacodynamics of the drugs, and safety concerns (5). Because each RNA-seq run generates tens to hundreds of millions of short reads with size ranging from 50bp-200bp, a tool that deciphers these short reads to an integrated and digestible analysis report is in high demand. QuickRNASeq (6) is an application for large-scale RNA-seq data analysis and real-time interactive visualization of complex data sets. This application automates the use of several of the best open-source tools to efficiently generate user friendly, easy to share, and ready to publish report. Figure 1 illustrates some of the interactive plots produced by QuickRNASeq. The visualization features of the application have been further improved since its first publication in early 2016. The original QuickRNASeq publication (6) provided details of background, software selection, and implementation. Here, we outline the steps required to implement QuickRNASeq in users own environment, as well as demonstrate some basic yet powerful utilities of the advanced interactive visualization modules in the report.\n\nO_FIG O_LINKSMALLFIG WIDTH=188 HEIGHT=200 SRC=\"FIGDIR/small/125856_fig1.gif\" ALT=\"Figure 1\">\nView larger version (59K):\norg.highwire.dtl.DTLVardef@1f6fb70org.highwire.dtl.DTLVardef@1f5a748org.highwire.dtl.DTLVardef@b990fborg.highwire.dtl.DTLVardef@dd5336_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFig. 1C_FLOATNO Interactive plots from QuickRNASeq report. Figures (a, b, c) can be retrieved by clicking on the pointing hands as shown in figure Id. On any of these interactive plots, mouse over each sample displays associated sample QC metrics, (a) Read mapping summary in the expanded display mode, (b) SNP concordance matrix of 48 samples from 5 donors. Samples from the same donor should be highly concordant, (c) Gene expression chart, which shows the number of genes past various expression thresholds, (d) Center portion of the QuickRNASeq report, (e) Parallel plot linking multiple QC measures for the same samples plus table of multi-dimensional QC measures.\n\nC_FIG

bioinformatics

LIN28 selectively modulates a subclass of let-7 microRNAs

LIN28 is a bipartite RNA-binding protein that post-transcriptionally inhibits let-7 microRNAs to regulate development and influence disease states. However, the mechanisms of let-7 suppression remains poorly understood, because LIN28 recognition depends on coordinated targeting by both the zinc knuckle domain (ZKD)--which binds a GGAG-like element in the precursor--and the cold shock domain (CSD), whose binding sites have not been systematically characterized. By leveraging single-nucleotide-resolution mapping of LIN28 binding sites in vivo, we determined that the CSD recognizes a (U)GAU motif. This motif partitions the let-7 family into Class I precursors with both CSD and ZKD binding sites and Class II precursors with ZKD but no CSD binding sites. LIN28 in vivo recognition--and subsequent 3' uridylation and degradation--of Class I precursors is more efficient, leading to their stronger suppression in LIN28-activated cells and cancers. Thus, CSD binding sites amplify the effects of the LIN28 activation with potential implication in development and cancer.

molecular biology

Genomic copy-number loss is rescued by self-limiting production of DNA circles

Copy-number changes generate phenotypic variability in health and disease. Whether organisms protect against copy-number changes is largely unknown. Here, we show that Saccharomyces cerevisiae monitors the copy number of its ribosomal DNA (rDNA) and rapidly responds to copy-number loss with the clonal amplification of extrachromosomal rDNA circles (ERCs) from chromosomal repeats. ERC production is proportional to repeat loss and reaches a dynamic steady state that responds to the addition of exogenous rDNA copies. ERC levels are also modulated by RNAPI activity and diet, suggesting that rDNA copy number is calibrated against the cellular demand for rRNA. Lastly, we show that ERCs reinsert into the genome in a dosage-dependent manner, indicating that they provide a reservoir for ultimately increasing rDNA array length. Our results reveal a DNA-based mechanism for rapidly restoring copy number in response to catastrophic gene loss that shares fundamental features with unscheduled copy-number amplifications in cancer cells.

molecular biology

cis-regulatory architecture of a short-range EGFR organizing center in the Drosophila melanogaster leg.

We characterized the establishment of an Epidermal Growth Factor Receptor (EGFR) organizing center (EOC) during leg development in Drosophila melanogaster. Initial EGFR activation occurs in the center of leg discs by expression of the EGFR ligand Vn and the EGFR ligand-processing protease Rho, each through single enhancers, vnE and rhoE, that integrate inputs from Wg, Dpp, Dll and Sp1. Deletion of vnE and rhoE eliminates vn and rho expression in the center of the leg imaginal discs, respectively. Animals with deletions of both vnE and rhoE (but not individually) show distal but not medial leg truncations, suggesting that the distal source of EGFR ligands acts at short-range to only specify distal-most fates, and that multiple additional ring enhancers are responsible for medial fates. Further, based on the cis-regulatory logic of vnE and rhoE we identified many additional leg enhancers, suggesting that this logic is broadly used by many genes during Drosophila limb development.\n\nAuthor SummaryThe EGFR signaling pathway plays a major role in innumerable developmental processes in all animals and its deregulation leads to different types of cancer, as well as many other developmental diseases in humans. Here we explored the integration of inputs from the Wnt- and TGF-beta signaling pathways and the leg-specifying transcription factors Distal-less and Sp1 at enhancer elements of EGFR ligands. These enhancers trigger a specific EGFR-dependent developmental output in the fly leg that is limited to specifying distal-most fates. Our findings suggest that activation of the EGFR pathway during fly leg development occurs through the activation of multiple EGFR ligand enhancers that are active at different positions along the proximo-distal axis. Similar enhancer elements are likely to control EGFR activation in humans as well. Such DNA elements might be hot spots that cause formation of EGFR-dependent tumors if mutations in them occur. Thus, understanding the molecular characteristics of such DNA elements could facilitate the detection and treatment of cancer.

developmental biology

Compressive Stress Enhances Invasive Phenotype of Cancer Cells via Piezo1 Activation

Uncontrolled growth in solid tumor generates compressive stress that drives cancer cells into invasive phenotypes, but little is known about how such stress affects the invasion and matrix degradation of cancer cells and the underlying mechanisms. Here we show that compressive stress enhanced invasion, matrix degradation, and invadopodia formation of breast cancer cells. We further identified Piezo1 channels as the putative mechanosensitive cellular components that transmit the compression to induce calcium influx, which in turn triggers activation of RhoA, Src, FAK, and ERK signaling, as well as MMP-9 expression. Interestingly, for the first time we observed invadopodia with matrix degradation ability on the apical side of the cells, similar to those commonly observed at the cells ventral side. Furthermore, we demonstrate that Piezo1 and caveolae were both involved in mediating the compressive stress-induced cancer cell invasive phenotype as Piezo1 and caveolae were often colocalized, and reduction of Cav-1 expression or disruption of caveolae with methyl-{beta}-cyclodextrin led to not only reduced Piezo1 expression but also attenuation of the invasive phenotypes promoted by compressive stress. Taken together, our data indicate that mechanical compressive stress activates Piezo1 channels to mediate enhanced cancer cell invasion and matrix degradation that may be a critical mechanotransduction pathway during, and potentially a novel therapeutic target for, breast cancer metastasis

cell biology

Learning common and specific patterns from data of multiple interrelated biological scenarios with matrix factorization

High-throughput biological technologies (e.g., ChIP-seq, RNA-seq and single-cell RNA-seq) rapidly accelerate the accumulation of genome-wide omics data in diverse interrelated biological scenarios (e.g., cells, tissues and conditions). Data dimension reduction and differential analysis are two common paradigms for exploring and analyzing such data. However, they are typically used in a separate or/and sequential manner. In this study, we propose a flexible non-negative matrix factorization framework CSMF to combine them into one paradigm to simultaneously reveal common and specific patterns from data generated under interrelated biological scenarios. We demonstrate the effectiveness of CSMF with four applications including pairwise ChIP-seq data describing the chromatin modification map on protein-DNA interactions between K562 and Huvec cell lines; pairwise RNA-seq data representing the expression profiles of two cancers (breast invasive carcinoma and uterine corpus endometrial carcinoma); RNA-seq data of three breast cancer subtypes; and single-cell sequencing data of human embryonic stem cells and differentiated cells at six time points. Extensive analysis yields novel insights into hidden combinatorial patterns embedded in these interrelated multi-modal data. Results demonstrate that CSMF is a powerful tool to uncover common and specific patterns with significant biological implications from data of interrelated biological scenarios.

bioinformatics

Inferring propensity amongst lung and breast carcinomas via overlapped gene expression profiles

Reconstruction of biological networks for topological analyses helps in correlation identification between various types of biomarkers. These networks have been vital components of System Biology in present era. Genes are the basic physical and structural unit of heredity. Genes act as instructions to make molecules called proteins. Alterations in the normal sequence of these genes are the root cause of various diseases and cancer is the prominent example disease caused by gene alteration or mutation. These slight alterations can be detected by microarray analysis. The high throughput data obtained by microarray experiments aid scientists in reconstructing cancer specific gene regulatory networks. The purpose of experiment performed is to find out the overlapping of the gene expression profiles of breast and lung cancer data, so that the common hub genes can be sifted and utilized as drug targets which could be used for the treatment of diseased conditions. In this study, first the differentially expressed genes have been identified (lung cancer and breast cancer), followed by a filtration approach and most significant genes are chosen using paired t-test and gene regulatory network construction. The obtained result has been checked and validated with the available databases and literature.

systems biology

Deep learning accurately predicts estrogen receptor status in breast cancer metabolomics data

Metabolomics holds the promise as a new technology to diagnose highly heterogeneous diseases. Conventionally, metabolomics data analysis for diagnosis is done using various statistical and machine learning based classification methods. However, it remains unknown if deep neural network, a class of increasingly popular machine learning methods, is suitable to classify metabolomics data. Here we use a cohort of 271 breast cancer tissues, 204 positive estrogen receptor (ER+) and 67 negative estrogen receptor (ER-), to test the accuracies of autoencoder, a deep learning (DL) framework, as well as six widely used machine learning models, namely Random Forest (RF), Support Vector Machines (SVM), Recursive Partitioning and Regression Trees (RPART), Linear Discriminant Analysis (LDA), Prediction Analysis for Microarrays (PAM), and Generalized Boosted Models (GBM). DL framework has the highest area under the curve (AUC) of 0.93 in classifying ER+/ER-patients, compared to the other six machine learning algorithms. Furthermore, the biological interpretation of the first hidden layer reveals eight commonly enriched significant metabolomics pathways (adjusted P-value<0.05) that cannot be discovered by other machine learning methods. Among them, protein digestion & absorption and ATP-binding cassette (ABC) transporters pathways are also confirmed in integrated analysis between metabolomics and gene expression data in these samples. In summary, deep learning method shows advantages for metabolomics based breast cancer ER status classification, with both the highest prediction accurcy (AUC=0.93) and better revelation of disease biology. We encourage the adoption of autoencoder based deep learning method in the metabolomics research community for classification.

bioinformatics

Hidden Markov Models Lead to Higher Resolution Maps of Mutation Signature Activity in Cancer

Knowing the activity of the mutational processes shaping a cancer genome may provide insight into tumorigenesis and personalized therapy. It is thus important to uncover the characteristic signatures of active mutational processes in patients from their patterns of single base substitutions. However, mutational processes do not act uniformly on the genome and are biased by factors such as the genomes chromatin structure or replication origins. These factors may lead to statistical dependencies among neighboring mutations, calling for modeling approaches that can account for such dependencies to better estimate mutational process activities.\n\nHere we develop the first sequence-dependent models for mutation signatures. We apply these models to characterize genomic and other factors that influence the activity of previously validated mutation signatures in breast cancer. We find that our tool, SO_SCPLOWIGC_SCPLOWMO_SCPLOWAC_SCPLOW, can accurately assign genomic mutations to mutation signatures, yielding assignments that are of higher likelihood than those obtained with models that assume independence between signatures and align better with current biological knowledge. Our analysis resolves a controversy related to the dependency of APOBEC signatures on replication time and links Signatures 18 and 30 to oxidative damage.\n\nModeling the sequential dependencies of mutation signatures leads to improved estimates of mutation signature activity both at the tumor-level and within specific genomic regions, yielding higher resolution maps of mutation signature activity in cancer.

bioinformatics

Systematic Functional Annotation and Visualization of Biological Networks

Large-scale biological networks map functional connections between most genes in the genome and can potentially uncover high level organizing principles governing cellular functions. These networks, however, are famously complex and often regarded as disordered masses of tangled interactions (\"hairballs\") that are nearly impenetrable to biologists. As a result, our current understanding of network functional organization is very limited. To address this problem, I developed a systematic quantitative approach for annotating biological networks and examining their functional structure. This method, named Spatial Analysis of Functional Enrichment (SAFE), detects network regions that are statistically overrepresented for a functional group or a quantitative phenotype of interest, and provides an intuitive visual representation of their relative positioning within the network. By successfully annotating the Saccharomyces cerevisiae genetic interaction network with Gene Ontology terms, SAFE proved to be sensitive to functional signals and robust to noise. In addition, SAFE annotated the network with chemical genomic data and uncovered a new potential mechanism of resistance to the anti-cancer drug bortezomib. Finally, SAFE showed that protein-protein interactions, despite their apparent complexity, also have a high level functional structure. These results demonstrate that SAFE is a powerful new tool for examining biological networks and advancing our understanding of the functional organization of the cell.

Systems Biology

Cytoplasmic volume and limiting nucleoplasmin scale nuclear size during Xenopus laevis development

How nuclear size is regulated relative to cell size is a fundamental cell biological question. Reductions in both cell and nuclear sizes during Xenopus laevis embryogenesis provide a robust scaling system to study mechanisms of nuclear size regulation. To test if the volume of embryonic cytoplasm is limiting for nuclear growth, we encapsulated gastrula stage embryonic cytoplasm and nuclei in droplets of defined volume using microfluidics. Nuclei grew and reached new steady-state sizes as a function of cytoplasmic volume, supporting a limiting component mechanism of nuclear size control. Through biochemical fractionation, we identified the histone chaperone nucleoplasmin (Npm2) as a putative nuclear size-scaling factor. Cellular amounts of Npm2 decrease over development, and nuclear size was sensitive to Npm2 levels both in vitro and in vivo, affecting nuclear histone levels and chromatin organization. Thus, reductions in cell volume with concomitant decreases in Npm2 amounts represent a developmental mechanism of nuclear size-scaling that may also be relevant to cancers with increased nuclear size.

cell biology

Pan-Cancer Analysis Reveals Technical Artifacts in The Cancer Genome Atlas (TCGA) Germline Variant Calls

The degree to which germline variation drives cancer development and shapes tumor phenotypes remains largely unexplored, possibly due to a lack of large scale publicly available germline data for a cancer cohort. Here we called germline variants on 9,618 cases from The Cancer Genome Atlas (TCGA) database representing 31 cancer types. We identified batch effects affecting loss of function (LOF) variant calls that can be traced back to differences in the way the sequence data were generated both within and across cancer types. Overall, LOF indel calls were more sensitive to technical artifacts than LOF Single Nucleotide Variant (SNV) calls. In particular, whole genome amplification of DNA prior to sequencing led to an artificially increased burden of LOF indel calls, which confounded association analyses relating germline variants to tumor type despite stringent indel filtering strategies. Due to the inherent noise we chose to remove all 614 amplified DNA samples, including all acute myeloid leukemia and virtually all ovarian cancer samples, from the final dataset. This study demonstrates how insufficient quality control can lead to false positive germlinetumor type associations and draws attention to the need to be sensitive to problems associated with a lack of uniformity in data generation in TCGA data.\n\nAuthor SummaryCancer research to date has largely focused on genetic aberrations specific to tumor tissue. In contrast, the degree to which germline, or inherited, variation contributes to tumorigenesis remains unclear, possibly due to a lack of accessible germline variant data. In this study we identify germline variants in 9,618 samples using raw germline exome data from The Cancer Genome Atlas (TCGA). There are substantial differences in the way exome sequence data was generated both across and within cancer types in TCGA. We observe that differences in sequence data generation introduced batch effects, or variation that is due to technical factors not true biological variation, in our variant data. Most notably, we observe that amplification of DNA prior to sequencing resulted in an excess of predicted damaging indel variants. We show how these batch effects can confound germline association analyses if not properly addressed. Our study highlights the difficulties of working with large public genomic datasets like TCGA where samples are collected over time and across data centers, and particularly cautions the use of amplified DNA samples for genetic association analyses.

genomics