Search bioRxiv⌕ Search

Biology subjects

Naznin, M. F. S.

Publications and source records attributed to Naznin, M. F. S..

2 recordsLinked to original sources

DDS-E-Sim: A Transformer-Based Generative Framework for Simulating Error-Prone Sequences in DNA Data Storage

DNA has emerged as a promising medium for long-lasting data stoage due to its high information density and long-term stability. However, DNA storage is a complex process where each stage introduces noise and errors. Since running DNA data storage experiments in vitro is still expensive and time-consuming, a simulation model is quite necessary that can mimic the error patterns in the real data and simulate the experiments. Existing tools often rely on fixed error rates or are specific to certain technologies. We propose DDS-E-Sim, a transformer-based probabilistic generative framework that simulates errors in a DNA data storage channel, regardless of the process or technology. DDS-E-Sim successfully captures the error distribution of DNA storage pipelines and learns to stochastically generate erroneous DNA reads. Given oligos (DNA sequences to write), it outputs erroneous reads resembling real pipelines capturing both random and biased errors, such as k-mer and transition errors. Evaluations on two distinct technology-specific datasets show high fidelity and universality: DDS-E-SIM exhibit a total error rate deviation of only 0.1% and 0.7% respectively on the datasets processed with Illumina MiSeq and Oxford Nanopore. Additionally, our simulator generates 100,743 unique oligos from 35,329 sequences, with coverage 5 (each sequence read five times) in the test datasets, demonstrating its ability to simulate biased errors and stochastic properties simultaneously.

bioinformatics↗

SEEDS: Simulating Emergence of Errors in DNA Storage

BackgroundDNA storage is a nonvolatile memory technology for storing data as synthetic DNA strings which offers unprecedented storage density and durability. Yet, the application of DNA as a practical digital information storage medium remains an enigma, since this is extremely expensive and it takes a substantial amount of time to encode and decode data to/from synthetic DNA. More importantly, various phases of DNA storage pipeline (e.g., synthesis, sequencing, etc.) are error prone. Furthermore, DNA is subject to decay over time and the reliability of the synthetic DNA depends on various aspects, including preservation medium and temperature. To allow for the perfect storage and recovery of the information and thereby making it competitive with the existing flash or tape based technologies, advanced error protection schemes are necessary. However, evaluating and comparing various DNA storage technologies and error correcting codes under realistic model conditions - comprising a wide array of synthesis medium, sequencing technologies, temperature and duration - is prohibitively time consuming and expensive. ResultsIn this study, we present SEEDS, an error model based simulator to mimic the process of accumulating errors at different phases of DNA storage. SEEDS is the first known simulator which incorporates various empirically derived statistical (or stochastic?) error models, mimicking the generation and propagation of different types of errors at various phases in DNA storage. It was assessed for its validity against the data from a number of published wet-lab experiments. ConclusionsSEEDS is easy to use and offers flexible and comprehensive parameter settings to mimic the error models in DNA storage. Validation against in vitro experimental results suggests its promise for emulating the stochastic models of error generation and propagation in DNA storage. SEEDS is available as a web interface with a server side application, along with portable cross-platform native applications (available at givethelink).

bioinformatics↗