Search bioRxivSearch

Biology subjects

Baudis, M.

Publications and source records attributed to Baudis, M..

4 recordsLinked to original sources

DNA copy number imbalances in primary cutaneous lymphomas (PCL)

Cutaneous lymphomas (CL) represent a clinically defined group of extranodal non-Hodgkin lymphomas harboring heterogeneous and incompletely delineated molecular aberrations. Over the past decades, molecular studies have identified several chromosomal aberrations, but the interpretation of individual genomic studies can be challenging.\n\nWe conducted a meta-analysis to delineate genomic alterations for different types of PCL. Searches of PubMed and ISI Web of Knowledge for the years 1996 to 2016 identified 32 publications reporting the investigation of PCL for genome-wide copy number alterations, by means of comparative genomic hybridization techniques and whole genome and exome sequencing. For 449 samples from 22 publications, copy number variation data was accessible for sample based meta-analysis. Summary profiles for genomic imbalances, generated from case-specific data, identified complex genomic imbalances, which could discriminate between different subtypes of CL and promise a more accurate classification. The collected data presented in this study are publicly available through the \"Progenetix\" online repository.

cancer biology

Population assignment from cancer genome profiling data

For a variety of human malignancies, incidence, treatment efficacy and overall prognosis show considerable variation between different populations and ethnic groups. Disentangling the effects related to particular population backgrounds can help in both understanding cancer biology and in tailoring therapeutic interventions. Because self-reported or inferred patient data can be incomplete or misleading due to migration and genomic admixture, a data-driven ancestry estimation should be preferred. While algorithms to analyze ancestry structure from healthy individuals have been developed, an easy-to-use tool to assign population groups based on genotyping data from SNP profiles is still missing and benchmarking for the validity of population assignment strategy for aberrant cancer genomes was not tested.\n\nWe benchmarked the consistency and accuracy of cross-platform population assignment. We also demonstrated its high accuracy to process unaltered as well as cancer genomes. Despite widespread and extensive somatic mutations of cancer profiling data, population assignment consistency between germline and highly mutated samples from cancer patients reached of 97% and 92% for assignment into 5 and 26 populations re-spectively. Comparison of our benchmarked results with self-reported meta-data estimated a matching rate between 88% to 92%. Despite a relatively high matching rate, the ethnicity labels indicated in meta-data are vague compared to the standardized output from our tool.\n\nWe have developed a bioinformatics tool to assign the populations from genome profiling data and validated its performance in healthy as well as aberrant cancer genomes. It is ready-to-use for genotyping data from nine commercial SNP array platforms or sequencing data. This tool is effective to scrutinize the population structure in cancer genomes and provides better measure to integrate genotyping data from various platforms instead of self-reported information. It will facilitate research on interplay between ethnicity related genetic background and molecular patterns in cancer entities and disentangling possible hereditary contributions.\n\nThe docker image of the tool is provided in DockerHub as \"baudisgroup/snp2pop\".

bioinformatics

A harmonized meta-knowledgebase of clinical interpretations of cancer genomic variants

Precision oncology relies on the accurate discovery and interpretation of genomic variants to enable individualized diagnosis, prognosis, and therapy selection. We found that knowledgebases containing clinical interpretations of somatic cancer variants are highly disparate in interpretation content, structure, and supporting primary literature, impeding consensus when evaluating variants and their relevance in a clinical setting. With the cooperation of experts of the Global Alliance for Genomics and Health (GA4GH) and six prominent cancer variant knowledgebases, we developed a framework for aggregating and harmonizing variant interpretations to produce a meta-knowledgebase of 12,856 aggregate interpretations covering 3,437 unique variants in 415 genes, 357 diseases, and 791 drugs. We demonstrated large gains in overlap between resources across variants, diseases, and drugs as a result of this harmonization. We subsequently demonstrated improved matching between a patient cohort and harmonized interpretations of potential clinical significance, observing an increase from an average of 33% per individual knowledgebase to 56% in aggregate. Our analyses illuminate the need for open, interoperable sharing of variant interpretation data. We also provide an open and freely available web interface (search.cancervariants.org) for exploring the harmonized interpretations from these six knowledgebases.

bioinformatics

segment_liftover: a Python tool to convert segments between genome assemblies

The process of assembling a species reference genome may be performed in a number of iterations, with subsequent genome assemblies differing in the coordinates of mapped elements. The conversion of genome coordinates between different assemblies is required for many integrative and comparative studies. While currently a number of bioinformatics tools are available to accomplish this task, most of them are tailored towards the conversion of single genome coordinates. When converting the boundary positions of segments spanning larger genome regions, segments may be mapped into smaller subsegments if the original segments continuity is disrupted in the target assembly. Such a conversion may lead to a relevant degree of data loss in some circumstances such as copy number variation (CNV) analysis, where the quantitative representation of a genomic region takes precedence over base-specific accuracy. segment_liftover aims at continuity-preserving remapping of genome segments between assemblies and provides features such as approximate locus conversion, automated batch processing and comprehensive logging to facilitate processing of datasets containing large numbers of structural genome variation data.

bioinformatics