The Tiling Algorithm - A general method for structural characterization of accurate long DNA sequence reads: application to AAV genome sequences.
Adeno-associated virus (AAV), a common vector used in human gene therapy, is challenging for DNA sequencing for several reasons. First, AAVs replication cycle results in structural rearrangements. Specifically, the inverted terminal repeat (ITR) at each end of the genome as well as the genetic payload can invert independently. Second, the ITR can prime DNA replication of the entire viral genome in the gap filling step of a sequencing library preparation protocol. Third, the AAV manufacturing process can produce small fractions of viral particles containing host cell DNA and/or fragments of the helper plasmids. And finally, short read sequencing is ill suited for the characterization of repetitive sequences and structural rearrangements involving several hundred nucleotides. Pacific Biosciences (PacBio) long-read sequencers are capable of full length, accurate, single molecule sequencing of AAV viral genomes, which addresses the challenge of working with repetitive sequences. However, sequence analysis methods based on alignment to a reference are confounded by the other three challenges above. We present a simple algorithm for determining the arrangement of functional elements of single DNA molecules which can be aggregated to provide a sensitive measure of the population of sequences in a sample, including minor species. Using data from four publicly available datasets, we demonstrate our algorithm is able to characterize nearly all of the species in the AAV samples.