Search bioRxiv⌕ Search

Biology subjects

Brunk, E. C.

Publications and source records attributed to Brunk, E. C..

2 recordsLinked to original sources

CytoCellDB: A Resource Database For Classification and Analysis of Extrachromosomal DNA in Cancer

Extrachromosomal DNA (ecDNA), or double minute chromosomes, are established cytogenetic markers for malignancy and genome instability. More recently, the cancer community has gained a heightened awareness of the roles of ecDNA in cancer proliferation, drug resistance and epigenetic remodeling. A current hindrance to understanding the biological roles of ecDNA is the lack of available cell line model systems with experimental cytogenetic data that confirm ecDNA status. Although several recent landmark studies have identified common cell lines and tumor models with ecDNA, the current sample size limits our ability to detect ecDNA-driven molecular differences due to limitations in power. Increasing the number of model systems known to express ecDNA would provide new avenues for understanding the fundamental underpinnings of ecDNA biology and would unlock a wealth of potential targeting strategies for ecDNA-driven cancers. To bridge this gap, we created CytoCellDB, a resource that provides karyotype annotations and leverages publicly available global cell line data from the Cancer Dependency Map (DepMap) and the Cancer Cell Line Encyclopedia (CCLE). Here, we identify 139 cell lines that express ecDNA, which is a 200% increase from the current sample size. We expanded the total number of cancer cell lines with ecDNA annotations to 577, which is a 400% increase or 31% of cell lines in CCLE/ DepMap. We demonstrate that a strength of CytoCellDB is the ability to interrogate ecDNA, and a compendium of other chromosomal aberrations, in the context of cancer-specific vulnerabilities, drug sensitivities, and molecular data (genomics, transcriptomics, methylation, proteomics). We anticipate that CytoCellDB will advance cytogenomics research and population-scale discoveries related to ecDNA as well as provide insights into strategies and best practices for determining novel therapeutics that overcome ecDNA-driven drug resistance.

genomics↗

Machine learning of three-dimensional protein structures to predict the functional impacts of genome variation

Research in the human genome sciences generates a substantial amount of genetic data for hundreds of thousands of individuals, which concomitantly increases the number of variants with unknown significance (VUS). Bioinformatic analyses can successfully reveal rare variants and variants with clear associations to disease-related phenotypes. These studies have made a significant impact on how clinical genetic screens are interpreted and how patients are stratified for treatment. There are few, if any, comparable computational methods for variants to biological activity predictions. To address this gap, we developed a machine learning method that uses protein three-dimensional structures from AlphaFold to predict how a variant will influence changes to a genes downstream biological pathways. We trained state-of-the-art machine learning classifiers to predict which protein regions will most likely impact transcriptional activities of two proto-oncogenes, nuclear factor erythroid 2 (NFE2)-related factor 2 (Nrf2) and c-MYC. We have identified classifiers that attain accuracies higher than 80%, which have allowed us to identify a set of key protein regions that lead to significant perturbations in c-MYC or Nrf2 transcriptional pathway activities. SignificanceThe vast majority of mutations are either unspecified and/or their downstream biological implications are poorly understood. We have created a method that utilizes protein structure to cluster mutations from population-scale repositories to predict downstream functional impacts. The broader impacts of this approach include advanced filtering of mutations that are likely to impact genome function.

bioinformatics↗