Search bioRxiv⌕ Search

Biology subjects

Klie, A.

Publications and source records attributed to Klie, A..

3 recordsLinked to original sources

Epigenetic Germline Variants Predict Cancer Prognosis and Risk and Distribute Uniquely in Topologically Associating Domains

Cancer is a highly heterogeneous disease caused by genetic and epigenetic alterations in normal cells. A recent study uncovered methylation quantitative trait loci (meQTLs) associated with different levels of local DNA methylation in cancers. Here, we investigated whether the distribution of cancer meQTLs reflected functional organization of the genome in the form of chromatin topologically associated domains (TADs), and evaluated whether cancer meQTLs near known driver genes have the potential to influence cancer risk or progression. At TAD boundaries, we observed differences in the distribution of meQTLs when one or both of the adjacent TADs was transcriptionally active, with higher densities near inactive TADs. Furthermore, we found differences in cancer meQTL distributions in active versus inactive TADs and observed an enrichment of meQTLs in active TADs near tumor suppressors, whereas there was a depletion of such meQTLs near oncogenes. Several meQTLs were associated with cancer risk in the UKBioBank, and we were able to reproduce breast cancer risk associations in the DRIVE cohort. Survival analysis in TCGA implicated a number of meQTLs in 13 tumor types. In 10 of these, polygenic meQTL scores were associated with increased hazard in a CoxPH analysis. Risk and survival-associated meQTLs tended to affect cancer genes involved in DNA damage repair and cellular adhesion and reproduced cancer-specific associations reported in prior literature. In summary, this study provides evidence that genetic variants that influence local DNA methylation are affected by chromatin structure and can impact tumor evolution.

bioinformatics↗

EUGENe: A Python toolkit for predictive analyses of regulatory sequences

Deep learning (DL) has become a popular tool to study cis-regulatory element function. Yet efforts to design software for DL analyses in genomics that are Findable, Accessible, Interoperable and Reusable (FAIR) have fallen short of fully meeting these criteria. Here we present EUGENe (Elucidating the Utility of Genomic Elements with Neural Nets), a FAIR toolkit for the analysis of labeled sets of nucleotide sequences with DL. EUGENe consists of a set of modules that empower users to execute the key functionality of a DL workflow: 1) extracting, transforming and loading sequence data from many common file formats, 2) instantiating, initializing and training diverse model architectures, and 3) evaluating and interpreting model behavior. We designed EUGENe to be simple; users can develop workflows on new or existing datasets with two customizable Python objects, annotated sequence data (SeqData) and PyTorch models (BaseModel). The modularity and simplicity of EUGENe also make it highly extensible and we illustrate these principles through application of the toolkit to three predictive modeling tasks. First, we train and compare a set of built-in models along with a custom architecture for the accurate prediction of activities of plant promoters from STARR-seq data. Next, we apply EUGENe to an RNA binding prediction task and showcase how seminal model architectures can be retrained in EUGENe or imported from Kipoi. Finally, we train models to classify transcription factor binding by wrapping functionality from Janngu, which can efficiently extract sequences in BED file format from the human genome. We emphasize that the code used in each use case is simple, readable, and well documented (https://eugene-tools.readthedocs.io/en/latest/index.html). We believe that EUGENe represents a springboard toward a collaborative ecosystem for DL applications in genomics research. EUGENe is available for download on GitHub (https://github.com/cartercompbio/EUGENe) along with several introductory tutorials and for installation on PyPi (https://pypi.org/project/eugene-tools/).

bioinformatics↗

Affinity-optimizing variants within cardiac enhancers disrupt heart development and contribute to cardiac traits

Enhancers direct precise gene expression patterns during development and harbor the majority of variants associated with disease. We find that suboptimal affinity ETS transcription factor binding sites are prevalent within Ciona and human developmental heart enhancers. Here we demonstrate in two diverse systems, Ciona intestinalis and human iPSC-derived cardiomyocytes (iPSC-CMs), that single nucleotide changes can optimize the affinity of ETS binding sites, leading to gain-of-function gene expression associated with heart phenotypes. In Ciona, ETS affinity-optimizing SNVs lead to ectopic expression and phenotypic changes including two beating hearts. In human iPSC-CMs, an affinity-optimizing SNV associated with QRS duration occurs within an SCN5A enhancer and leads to increased enhancer activity. Our mechanistic approach provides a much-needed systematic framework that works across different enhancers, cell types and species to pinpoint causal enhancer variants contributing to enhanceropathies, phenotypic diversity and evolutionary changes. In BriefThe prevalent use of low-affinity ETS sites within developmental heart enhancers creates vulnerability within genomes whereby single nucleotide changes can dramatically increase binding affinity, causing gain-of-function enhancer activity that impacts heart development. HighlightsETS affinity-optimizing SNVs can lead to migration defects and a multi-chambered heart. An ETS affinity-optimizing human SNV within an SCN5A enhancer increases expression and is associated with QRS duration. Searching for ETS affinity-optimizing variants is a systematic and generalizable approach to pinpoint causal enhancer variants.

genomics↗