Search bioRxiv⌕ Search

Biology subjects

Gandolfi, G.

Publications and source records attributed to Gandolfi, G..

2 recordsLinked to original sources

The Minimal Dataset for Cancer of the 1+Million Genomes Initiative

For a real impact on healthcare, precision cancer medicine requires accessibility and interoperability of clinical and genomic data across centres and countries. Due to the heterogeneous digitization in Europe and worldwide, the definition of models for standardised data collection and usability becomes mandatory if countries want to work together on this mission. The European Union 1+Million Genomes (1+MG) initiative, supported by the Horizon 2020 Beyond 1 Million Genomes project, aims at outlining data models, guidance, best practices, and technical infrastructures for transnational access to sequenced genomes, including cancer genomes. Within the framework of the cancer-focused Working Group 9, we developed the 1+MG-Minimal Dataset for Cancer (1+MG-MDC)-a data model encompassing 140 items and organized in eight conceptual domains for the collection of cancer-related clinical information and genomics metadata. The 1+MG-MDC, which results from a multidisciplinary effort, leverages pre-existing models and emphasizes the annotation and traceability of multiple aspects relevant to the complex longitudinal path of the cancer disease and its treatment. We strived to make the 1+MG-MDC easy to adopt, yet comprehensive, addressing the needs of both clinicians and researchers. We will periodically revise and update it to ensure it remains fit for purpose. We propose the 1+MG-MDC as a model to create homogeneous databases, which would, in turn, guide discussions on clinical and genomic features with prognostic or therapeutic value and foster real-world data research.

cancer biology↗

Scalable Integration of Multiomic Single Cell Data Using Generative Adversarial Networks

Single cell profiling has become a common practice to investigate the complexity of tissues, organs and organisms. Recent technological advances are expanding our capabilities to profile various molecular layers beyond the transcriptome such as, but not limited to, the genome, the epigenome and the proteome. Depending on the experimental procedure, these data can be obtained from separate assays or from the very same cells. Despite development of computational methods for data integration is an active research field, most of the available strategies have been devised for the joint analysis of two modalities and cannot accommodate a high number of them. To solve this problem, we here propose a multiomic data integration framework based on Wasserstein Generative Adversarial Networks (MOWGAN) suitable for the analysis of paired or unpaired data with high number of modalities (>2). At the core of our strategy is a single network trained on all modalities together, limiting the computational burden when many molecular layers are evaluated. Source code of our framework is available at https://github.com/vgiansanti/MOWGAN.

bioinformatics↗