Search bioRxivSearch

Biology subjects

Eils, R.

Publications and source records attributed to Eils, R..

16 recordsLinked to original sources

Impact of cancer mutational signatures on transcription factor motifs in the human genome

BackgroundSomatic mutations in cancer genomes occur through a variety of molecular mechanisms, which contribute to different mutational patterns. To summarize these, mutational signatures have been defined using a large number of cancer genomes, and related to distinct mutagenic processes. Each cancer genome can be compared to this reference dataset and its exposure to one or the other signature be determined. Given the very different mutational patterns of these signatures, we anticipate that they will have distinct impact on genomic elements, in particular motifs for transcription factor binding sites (TFBS).\n\nResultsIn this work, we build the link between mutational signatures and TFBS motif alterations. We investigated and computed the theoretical impact of mutational signatures on 512 TFBS motifs, hence translating the trinucleotide mutation frequencies of the signatures into alteration frequencies of specific TFBS motifs, leading either to creation of disruption of these motifs. We further build a theoretical prediction of the alteration patterns for different cancer types based on the exposure of these cancer types to the mutation signatures. For certain motifs, a high correlation is observed between the TFBS motif creation and disruption events related to the information content of the motif.\n\nConclusionOur results show that the mutational signatures have different impact on the binding motifs of transcription factors and that for certain high complexity motifs there is a strong correlation between creation and disruption, related to the information content of the motif. This study represents a background estimation of the alterations due purely to mutational signatures in the absence of additional contributions, e.g. from evolutionary processes.

bioinformatics

HTSlib-js: Developer toolkit for a client-side JavaScript interface for HTSlib

BackgroundAn increasing number of bioinformatics tools are developed in JavaScript to provide an interactive visual interface to users. However, there are no tools yet that enable fast client-side analysis of large, high-throughput sequencing file formats, such as BAM and VCF files.\n\nResultsWe present HTSlib-js, an asm.js-based JavaScript wrapper for HTSlib, the de facto standard for processing BAM and VCF files. HTSlib-js exploits recent technological advances in web browser engines that dramatically increase the performance of browser-based tools to enable swift processing of files in the aforementioned formats. HTSlib-js enables quick development of JavaScript-based applications that include processing of aligned sequence reads and variant calling data.\n\nConclusionsHTSlib-js constitutes a toolkit for developers to easily write fast, accessible, browser-based applications that place an emphasis on data visualization, interactivity and privacy. Real-world examples demonstrate the capabilities of HTSlib-js and serve as guide for own developments.

bioinformatics

Web-based design and analysis tools for CRISPR base editing

BackgroundAs a result of its simplicity and high efficiency, the CRISPR-Cas system has been widely used as a genome editing tool. Recently, CRISPR base editors, which consist of deactivated Cas9 (dCas9) or Cas9 nickase (nCas9) linked with a cytidine or a guanine deaminase, have been developed. Base editing tools will be very useful for gene correction because they can produce highly specific DNA substitutions without the introduction of any donor DNA, but dedicated web-based tools to facilitate the use of such tools have not yet been developed.\n\nResultsWe present two web tools for base editors, named BE-Designer and BE-Analyzer. BE-Designer provides all possible base editor target sequences in a given input DNA sequence with useful information including potential off-target sites. BE-Analyzer, a tool for assessing base editing outcomes from next generation sequencing (NGS) data, provides information about mutations in a table and interactive graphs. Furthermore, because the tool runs client-side, large amounts of targeted deep sequencing data (>100MB) do not need to be uploaded to a server, substantially reducing running time and increasing data security. BE-Designer and BE-Analyzer can be freely accessed at http://www.rgenome.net/bedesigner/ and http://www.rgenome.net/be-analyzer/respectively\n\nConclusionWe develop two useful web tools to design target sequence (BE-Designer) and to analyze NGS data from experimental results (BE-Analyzer) for CRISPR base editors.

bioinformatics

Single-fluorescent protein reporters allow parallel quantification of NK cell-mediated granzyme and caspase activities in single target cells

1.Natural killer (NK) cells eliminate infected and tumorigenic cells through delivery of granzymes via perforin pores or by activation of caspases via death receptors. In order to understand how NK cells combine different cell death mechanisms it is important to quantify target cell responses on a single cell level. However, currently existing reporters do not allow the measurement of several protease activities inside the same cell. Here we present a strategy for the comparison of two different proteases at a time inside individual target cells upon engagement by NK cells. We developed single-fluorescent protein reporters containing the RIEAD or the VGPD cleavage site for the measurement of granzyme B activity. We show that these two granzyme B reporters can be applied in combination with caspase-8 or caspase-3 reporters. While we did not find that caspase-8 was activated by granzyme B, our method revealed that caspase-3 activity follows granzyme B activity with a delay of about 6 minutes. Finally, we illustrate the comparison of several different reporters for granzyme A, M, K and H. The here presented approach is a valuable means for the investigation of the temporal evolution of cell death mediated by cytotoxic lymphocytes.

immunology

Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity

Integrative analysis of multi-omics layers at single cell level is critical for accurate dissection of cell-to-cell variation within certain cell populations. Here we report scCAT-seq, a technique for simultaneously assaying chromatin accessibility and the transcriptome within the same single cell. We show that the combined single cell signatures enable accurate construction of regulatory relationships between cis-regulatory elements and the target genes at single-cell resolution, providing a new dimension of features that helps direct discovery of regulatory patterns specific to distinct cell identities. Moreover, we generated the first single cell integrated maps of chromatin accessibility and transcriptome in human pre-implantation embryos and demonstrated the robustness of scCAT-seq in the precise dissection of master transcription factors in cells of distinct states during embryo development. The ability to obtain these two layers of omics data will help provide more accurate definitions of \"single cell state\" and enable the deconvolution of regulatory heterogeneity from complex cell populations.

genomics

pheno-seq - linking 3D phenotypes of clonal tumor spheroids to gene expression

3D-culture systems have advanced cancer modeling by reflecting physiological characteristics of in-vivo tissues, but our understanding of functional intratumor heterogeneity including visual phenotypes and underlying gene expression is still limited. Single-cell RNA-sequencing is the method of choice to dissect transcriptional tumor cell heterogeneity in an unbiased way, but this approach is limited in correlating gene expression with contextual cellular phenotypes.\n\nTo link morphological features and gene expression in 3D-culture systems, we present pheno-seq for integrated high-throughput imaging and transcriptomic profiling of clonal tumor spheroids. Specifically, we identify characteristic EMT expression signatures that are associated with invasive growth behavior in a 3D breast cancer model. Additionally, pheno-seq determined transcriptional programs containing lineage-specific markers that can be linked to heterogeneous proliferative capacity in a patient-derived 3D model of colorectal cancer. Finally, we provide evidence that pheno-seq identifies morphology-specific genes that are missed by scRNA-seq and inferred single-cell regulatory states without acquiring additional single cell expression profiles. We anticipate that directly linking molecular features with patho-phenotypes of cancer cells will improve the understanding of intratumor heterogeneity and consequently be useful for translational research.

genomics

CD95 receptor activation by ligand-induced trimerization is independent of its partial pre-ligand assembly

CD95 (Fas, APO-1, TNFRSF6) is a widely expressed single-pass transmembrane protein that is implicated in cell death, inflammatory response, proliferation and cell migration. CD95 ligand (CD95L, FasL, TNFSF6), is a potent apoptotic inducer in the membrane form but not when cleaved into soluble CD95L (sCD95L). Here, we aimed at understanding the relation between ligand-receptor multimerization and receptor activation by correlating the kinetics of ligand binding, receptor oligomerization, FADD (FAS-Associated via Death Domain) recruitment and caspase-8 activation inside living cells. Using single molecule localization microscopy and Forster resonance energy transfer imaging we show that the majority of CD95 receptors on the plasma membrane are monomeric at rest. This was confirmed functionally as the wild-type receptor is not blocked by a receptor mutant that cannot bind ligand. Moreover, using time-resolved fluorescence imaging approaches we demonstrated that receptor multimerization follows instantaneously ligand binding, whereas FADD recruitment is delayed. This process can explain the typical delay time seen with caspase-8 activity reporters. Finally, the low activity of sCD95L, which was caused by inefficient FADD recruitment, was not explained by the low avidity for the receptor but by a receptor clustering mechanism that was different from the one induced by the strong apoptosis inducer IZ-sCD95L. Our results reveal that receptor activation is modulated by the capacity of its ligand to trimerize it.\n\nHighlightsO_LIAt a density of less than 10 receptors per {micro}m2 CD95 exists as monomer (58%) and dimer (42%)\nC_LIO_LIPre-formed dimers do not contribute to ligand-induced CD95 apoptotic signaling\nC_LIO_LIThe PLAD of CD95 attenuates overexpression-induced, ligand-independent cell death\nC_LIO_LIsoluble CD95L can rapidly multimerize CD95 after binding but it is still a poor inducer of apoptosis through inefficient FADD recruitment\nC_LIO_LIFADD recruitment kinetics but not ligand binding kinetics correlates with caspase-8 onset of activity\nC_LI

cell biology

Unraveling mitotic protein networks by 3D multiplexed epitope drug screening

Three-dimensional protein localization intricately determines the functional coordination of cellular processes. The complex spatial context of protein landscape has been assessed by multiplexed immunofluorescent staining1-3 or mass spectrometry4, applied to 2D cell culture with limited physiological relevance5 or tissue sections. Here, we present 3D SPECS, an automated technology for 3D Spatial characterization of Protein Expression Changes by microscopic Screening. This workflow encompasses iterative antibody staining of proteins, high-content imaging, and machine learning based classification of mitotic states. This is followed by mapping of spatial protein localization into a spherical, cellular coordinate system, the basis used for model-based prediction of spatially resolved affinities of various mitotic proteins. As a proof-of-concept, we mapped twelve epitopes in 3D cultured epithelial breast spheroids and investigated the network effects of mitotic cancer drugs with known limited success in clinical trials6-8. Our approach reveals novel insights into spindle fragility and global chromatin stress, and predicts unknown interactions between proteins in specific mitotic pathways. 3D SPECSs ability to map potential drug targets by multiplexed immunofluorescence in 3D cell cultured models combined with our automized high content assay will inspire future functional protein expression and drug assays.

systems biology

Identification and prioritisation of causal variants in human genetic disorders from exome or whole genome sequencing data

With genome sequencing entering the clinics as diagnostic tool to study genetic disorders, there is an increasing need for bioinformatics solutions that enable precise causal variant identification in a timely manner.\n\nBackgroundWorkflows for the identification of candidate disease-causing variants perform usually the following tasks: i) identification of variants; ii) filtering of variants to remove polymorphisms and technical artifacts; and iii) prioritization of the remaining variants to provide a small set of candidates for further analysis.\n\nMethodsHere, we present a pipeline designed to identify variants and prioritize the variants and genes from trio sequencing or pedigree-based sequencing data into different tiers.\n\nResultsWe show how this pipeline was applied in a study of patients with neurodevelopmental disorders of unknown cause, where it helped to identify the causal variants in more than 35% of the cases.\n\nConclusionsClassification and prioritization of variants into different tiers helps to select a small set of variants for downstream analysis.

bioinformatics

ACEseq - allele specific copy number estimation from whole genome sequencing

ACEseq is a computational tool for allele-specific copy number estimation in tumor genomes based on whole genome sequencing. In contrast to other tools it features GC-bias correction, unique replication timing-bias correction and integration of structural variant (SV) breakpoints for improved genome segmentation. ACEseq clearly outperforms widely used state-of-the art methods, provides a fully automated estimation of tumor cell content and ploidy, and additionally computes homologous recombination deficiency scores.

bioinformatics

Deciphering programs of transcriptional regulation by combined deconvolution of multiple omics layers

Metazoans are crucially dependent on multiple layers of gene regulatory mechanisms which allow them to control gene expression across developmental stages, tissues and cell types. Multiple recent research consortia have aimed to generate comprehensive datasets to profile the activity of these cell type- and condition-specific regulatory landscapes across many different cell lines and primary cells. However, extraction of genes or regulatory elements specific to certain entities from these datasets remains challenging. We here propose a novel method based on non-negative matrix factorization for disentangling and associating huge multi-assay datasets including chromatin accessibility and gene expression data. Taking advantage of implementations of NMF algorithms in the GPU CUDA environment full datasets composed of tens of thousands of genes as well as hundreds of samples can be processed without the need for prior feature selection to reduce the input size. Applying this framework to multiple layers of genomic data derived from human blood cells we unravel mechanisms of regulation of cell type-specific expression in T-cells and monocytes.

genomics

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics

Genomic footprints of activated telomere maintenance mechanisms in cancer

Cancers require telomere maintenance mechanisms for unlimited replicative potential. We dissected whole-genome sequencing data of over 2,500 matched tumor-control samples from 36 different tumor types to characterize the genomic footprints of these mechanisms. While the telomere content of tumors with ATRX or DAXX mutations (ATRX/DAXXtrunc) was increased, tumors with TERT modifications showed a moderate decrease of telomere content. One quarter of all tumor samples contained somatic integrations of telomeric sequences into non-telomeric DNA. With 80% prevalence, ATRX/DAXXtrunc tumors display a 3-fold enrichment of telomere insertions. A systematic analysis of telomere composition identified aberrant telomere variant repeat (TVR) distribution as a genomic marker of ATRX/DAXXtrunc tumors. In this clinically relevant subgroup, singleton TTCGGG and TTTGGG TVRs (previously undescribed) were significantly enriched or depleted, respectively. Overall, our findings provide new insight into the recurrent genomic alterations that are associated with the establishment of different telomere maintenance mechanisms in cancer.

genomics

Framework For Quality Assessment Of Whole Genome, Cancer Sequences

Working with cancer whole genomes sequenced over a period of many years in different sequencing centres requires a validated framework to compare the quality of these sequences. The Pan-Cancer Analysis of Whole Genomes (PCAWG) of the International Cancer Genome Consortium (ICGC), a project a cohort of over 2800 donors provided us with the challenge of assessing the quality of the genome sequences. A non-redundant set of five quality control (QC) measurements were assembled and used to establish a star rating system. These QC measures reflect known differences in sequencing protocol and provide a guide to downstream analyses of these whole genome sequences. The resulting QC measures also allowed for exclusion samples of poor quality, providing researchers within PCAWG, and when the data is released for other researchers, a good idea of the sequencing quality. For a researcher wishing to apply the QC measures for their data we provide a Docker Container of the software used to calculate them. We believe that this is an effective framework of quality measures for whole genome, cancer sequences, which will be a useful addition to analytical pipelines, as it has to the PCAWG project.

genomics

The Human Cell Atlas

The recent advent of methods for high-throughput single-cell molecular profiling has catalyzed a growing sense in the scientific community that the time is ripe to complete the 150-year-old effort to identify all cell types in the human body, by undertaking a Human Cell Atlas Project as an international collaborative effort. The aim would be to define all human cell types in terms of distinctive molecular profiles (e.g., gene expression) and connect this information with classical cellular descriptions (e.g., location and morphology). A comprehensive reference map of the molecular state of cells in healthy human tissues would propel the systematic study of physiological states, developmental trajectories, regulatory circuitry and interactions of cells, as well as provide a framework for understanding cellular dysregulation in human disease. Here we describe the idea, its potential utility, early proofs-of-concept, and some design considerations for the Human Cell Atlas.

cell biology

Resolving Drug Effects In Patient-Derived Cancer Cells Links Organoid Responses To Genome Alterations

Cancer drug screening in patient-derived cells holds great promise for personalized oncology and drug discovery but lacks standardization. Whether cells are cultured as conventional monolayer or advanced organoid cultures influences drug effects and thereby drug selection and clinical success. To precisely compare drug profiles in differently cultured primary cells, we developed DeathPro, an automated microscopy-based assay to resolve drug-induced cell death and proliferation inhibition. Using DeathPro, we screened cells from ovarian cancer patients in monolayer or organoid culture with clinically relevant drugs. Drug-induced growth arrest and efficacy of cytostatic drugs differed between the two culture systems. Interestingly, drug effects in organoids were more diverse and had lower therapeutic potential. Genomic analysis revealed novel links between drug sensitivity and DNA repair deficiency in organoids that were undetectable in monolayers. Thus, our results highlight the dependency of cytostatic drugs and pharmacogenomic associations on culture systems, and guide culture selection for drug tests.

cancer biology