Search bioRxivSearch

Biology subjects

de Ridder, J.

Publications and source records attributed to de Ridder, J..

5 recordsLinked to original sources

Molecular Heterogeneity and Early Metastatic Clone Selection in Testicular Germ Cell Cancer Development

BackgroundTesticular germ cell cancer (TGCC), being the most frequent malignancy in young Caucasian males, is initiated from an embryonic germ cell. This study determines intratumor heterogeneity to unravel tumor progression from initiation till metastasis.\n\nMethodsIn total 42 purified samples of four treatment-resistant nonseminomatous TGCC (NS) were investigated, including the precursor germ cell neoplasia in situ (GCNIS) and metastatic specimens, using whole genome- and targeted sequencing. Their evolution was reconstructed.\n\nResultsIntratumor molecular heterogeneity did not correspond to the supposed primary tumor histological evolution. Metastases after systemic treatment could be derived from cancer stem cells not identified in the primary cancer. GCNIS mostly lacked the molecular marks of the primary NS and comprised dominant clones that failed to progress. A BRCA-like mutational signature was observed without evidence for direct involvement of BRCA1 and BRCA2 genes.\n\nConclusionsOur data strongly support the hypothesis that NS is initiated by whole genome duplication, followed by chromosome copy number alterations in the cancer stem cell population, and accumulation of low numbers of somatic mutations. These observations of heterogeneity at all stages of tumorigenesis should be considered when treating patients with GCNIS-only disease, or with clinically overt NS.

cancer biology

A data-driven interactome of synergistic genes improves network based cancer outcome prediction

Robustly predicting outcome for cancer patients from gene expression is an important challenge on the road to better personalized treatment. Network-based outcome predictors (NOPs), which considers the cellular wiring diagram in the classification, hold much promise to improve performance, stability and interpretability of identified marker genes. Problematically, reports on the efficacy of NOPs are conflicting and for instance suggest that utilizing random networks performs on par to networks that describe biologically relevant interactions. In this paper we turn the prediction problem around: instead of using a given biological network in the NOP, we aim to identify the network of genes that truly improves outcome prediction. To this end, we propose SyNet, a gene network constructed ab initio from synergistic gene pairs derived from survival-labelled gene expression data. To obtain SyNet, we evaluate synergy for all 69 million pairwise combinations of genes resulting in a network that is specific to the dataset and phenotype under study and can be used to in a NOP model. We evaluated SyNet and 11 other networks on a compendium dataset of >4000 survival-labelled breast cancer samples. For this purpose, we used cross-study validation which more closely emulates real world application of these outcome predictors. We find that SyNet is the only network that truly improves performance, stability and interpretability in several existing NOPs. We show that SyNet overlaps significantly with existing gene networks, and can be confidently predicted (~85% AUC) from graph-topological descriptions of these networks, in particular the breast tissue-specific network. Due to its data-driven nature, SyNet is not biased to well-studied genes and thus facilitates post-hoc interpretation. We find that SyNet is highly enriched for known breast cancer genes and genes related to e.g. histological grade and tamoxifen resistance, suggestive of a role in determining breast cancer outcome.\n\nAuthor SummaryCancer is caused by disrupted activity of several pathways. Therefore, outcome predictors analyze patients expression profiles from perspective of gene groups collected from interactomes (e.g. protein interaction networks). These Network based Outcome Predictors (NOPs) hold potential to facilitate identification of dysregulated pathways and delivering improved prognosis. Nonetheless, recent studies revealed that compared to classical models, neither performance nor consistency can be improved using NOPs.\n\nWe argue that NOPs can only perform well under guidance of suitable networks. The commonly used networks may miss associations specially for under-studied genes. Additionally, these networks are often generic with low resemblance to perturbations that arise in cancer.\n\nTo address this issue, we exploit ~4100 samples and infer a disease specific network called SyNet linking synergistic gene pairs that collectively show predictivity beyond individual performance of genes.\n\nUsing identical datasets, we show that a NOP yields superior performance merely by considering groups of genes in SyNet. Further, NOP performance severely reduces if SyNet nodes are shuffled, confirming relevance of SyNet links.\n\nDue to simplicity of our approach, this framework can be used for any phenotype of interest. Our findings represent the value of network-based models and crucial role of interactome in their performance.

bioinformatics

Mining the forest: uncovering biological mechanisms by interpreting Random Forests

Biological datasets are large and complex. Machine learning models are therefore essential to capture relationships in the data. Unfortunately, the inferred complex models are often difficult to understand and interpretation is limited to a list of features ranked on their importance in the model.\n\nWe propose a computational approach, called Foresight, that enables interpretation of the patterns uncovered by Random Forest models trained on biological datasets. Foresight exploits the correlation structure in the data to uncover relevant groups of features and the interactions between them. This facilitates interpretation of the computational model and can provide more detailed insight in the underlying biological relationships than simply ranking features. We demonstrate Foresight on both an artificial dataset and a large gene expression dataset of breast cancer patients. Using the latter dataset we show that our approach retrieves biologically relevant features and provides a rich description of the interactions and correlation structure between these features.

bioinformatics

Locus-Specific Enhancer Hubs And Architectural Loop Collisions Uncovered From Single Allele DNA Topologies

Chromatin folding is increasingly recognized as a regulator of genomic processes such as gene activity. Chromosome conformation capture (3C) methods have been developed to unravel genome topology through the analysis of pair-wise chromatin contacts and have identified many genes and regulatory sequences that, in populations of cells, are engaged in multiple DNA interactions. However, pair-wise methods cannot discern whether contacts occur simultaneously or in competition on the individual chromosome. We present a novel 3C method, Multi-Contact 4C (MC-4C), that applies Nanopore sequencing to study multi-way DNA conformations of tens of thousands individual alleles for distinction between cooperative, random and competing interactions. MC-4C can uncover previously missed structures in sub-populations of cells. It reveals unanticipated cooperative clustering between regulatory chromatin loops, anchored by enhancers and gene promoters, and CTCF and cohesin-bound architectural loops. For example, we show that the constituents of the active b-globin super-enhancer cooperatively form an enhancer hub that can host two genes at a time. We also find cooperative interactions between further dispersed regulatory sequences of the active proto-cadherin locus. When applied to CTCF-bound domain boundaries, we find evidence that chromatin loops can collide, a process that is negatively regulated by the cohesin release factor WAPL. Loop collision is further pronounced in WAPL knockout cells, suggestive of a \"cohesin traffic jam\". In summary, single molecule multi-contact analysis methods can reveal how the myriad of regulatory sequences spatially coordinate their actions on individual chromosomes. Insight into these single allele higher-order topological features will facilitate interpreting the consequences of natural and induced genetic variation and help uncovering the mechanisms shaping our genome.

genetics

Mapping And Phasing Of Structural Variation In Patient Genomes Using Nanopore Sequencing

Structural genomic variants form a common type of genetic alteration underlying human genetic disease and phenotypic variation. Despite major improvements in genome sequencing technology and data analysis, the detection of structural variants still poses challenges, particularly when variants are of high complexity. Emerging long-read single-molecule sequencing technologies provide new opportunities for detection of structural variants. Here, we demonstrate sequencing of the genomes of two patients with congenital abnormalities using the ONT MinION at 11x and 16x mean coverage, respectively. We developed a bioinformatic pipeline - NanoSV - to efficiently map genomic structural variants (SVs) from the long-read data. We demonstrate that the nanopore data are superior to corresponding short-read data with regard to detection of de novo rearrangements originating from complex chromothripsis events in the patients. Additionally, genome-wide surveillance of SVs, revealed 3,253 (33%) novel variants that were missed in short-read data of the same sample, the majority of which are duplications < 200bp in size. Long sequencing reads enabled efficient phasing of genetic variations, allowing the construction of genome-wide maps of phased SVs and SNVs. We employed read-based phasing to show that all de novo chromothripsis breakpoints occurred on paternal chromosomes and we resolved the long-range structure of the chromothripsis. This work demonstrates the value of long-read sequencing for screening whole genomes of patients for complex structural variants.

genomics