Search bioRxivSearch

Biology subjects

Shortt, J. A.

Publications and source records attributed to Shortt, J. A..

2 recordsLinked to original sources

Finding and extending ancient simple sequence repeat-derived regions in the human genome

BackgroundPreviously, 3% of the human genome has been annotated as simple sequence repeats (SSRs), similar to the proportion annotated as protein coding. The origin of much of the genome is not well annotated, however, and some of the unidentified regions are likely to be ancient SSR-derived regions not identified by current methods. The identification of these regions is complicated because SSRs appear to evolve through complex cycles of expansion and contraction, often interrupted by mutations that alter both the repeated motif and mutation rate. We applied an empirical, kmer-based, approach to identify genome regions that are likely derived from SSRs.\n\nResultsThe sequences flanking annotated SSRs are enriched for similar sequences and for SSRs with similar motifs, suggesting that the evolutionary remains of SSR activity abound in regions near obvious SSRs. Using our previously described P-clouds approach, we identified SSR-clouds, groups of similar kmers (or oligos) that are enriched near a training set of unbroken SSR loci, and then used the SSR-clouds to detect likely SSR-derived regions throughout the genome.\n\nConclusionsOur analysis indicates that the amount of likely SSR-derived sequence in the human genome is 6.77%, over twice as much as previous estimates, including millions of newly identified ancient SSR-derived loci. SSR-clouds identified poly-A sequences adjacent to transposable element termini in over 74% of the oldest class of Alu (roughly, AluJ), validating the sensitivity of the approach. Poly-As annotated by SSR-clouds also had a length distribution that was more consistent with their poly-A origins, with mean about 35 bp even in older Alus. This work demonstrate that the high sensitivity provided by SSR-Clouds improves the detection of SSR-derived regions and will enable deeper analysis of how decaying repeats contribute to genome structure.

genomics

SNAPPY: Single Nucleotide Assignment of Phylogenetic Parameters on the Y chromosome

SummaryThe assignment of Y chromosome data to related clusters, or haplogroups, is a common application in human population genetics. To enable this at scale, we developed SNAPPY. SNAPPY is a software program used to assign Y-chromosome phylogeny-informed haplotypes using dense genotype data. The program efficiently tests all haplotypes in a provided Y-chromosome database to find the haplogroup that is best supported by the input genotypes. Importantly, the method considers both the amount of support for the specific haplogroup, as well as its ancestral haplogroups via parsimony. This accounts for the underlying genealogy the haplotypes represent, strengthening the accuracy of the assignments. SNAPPY is fast, scalable, and uses standard file formats, making it easy to integrate into analytical pipelines.\n\nAvailability and ImplementationThe program is implemented in python. The program, a user manual, haplotype databases, and test datasets are available for download at github.com/chrisgene/snappy.\n\nContactJonathan.shortt@ucdenver.edu, Chris.gignoux@ucdenver.edu

bioinformatics