A Comparison of Computational Methods for Modeling Stochastic Collaborative DNA Methylation Dynamics
DNA methylation is a widespread epigenetic modification that is important in gene regulation, cancer, and aging. Methylation occurs primarily on symmetric cytosine-phosphate-guanine (CpG) dinucleotides in vertebrates. Recent findings support that the enzymatic reactions governing both addition and removal of DNA methylation at an individual CpG site are influenced by methylation levels at other CpG sites in the vicinity. Mathematical models that treat this phenomenon have been termed "collaborative" models. These models are computationally challenging due to rare switching events and the combinatorial explosion of site-to-site interactions, hindering integration with data. We compare the efficiency and accuracy of three collaborative DNA methylation modeling approaches: a stochastic simulation algorithm (SSA), an exact Chemical Master Equation (CME), and a mean-field CME. The exact CME model is limited to very small system sizes (few CpGs), with computational expense scaling as 32N for a system of N sites. The SSA approach can accurately handle larger system sizes, with computational expense scaling as N2. The mean-field CME model accurately captures qualitative methylation patterns and is efficient for small (N<100) system sizes, but accuracy and efficiency suffer at larger system sizes, with expense scaling as N5.5. Using the developed numerical approaches, we compute methylation phase diagrams, which reveal how methylation levels and bimodality depend on the DNA-sequence-architecture of CpG clusters, as well as on enzymatic model parameters. We also present phase diagrams derived from the mouse epigenome. These data reveal a complex dependence of cluster methylation patterns on cluster architecture, which is nevertheless well-captured by the relatively simplified model.