Search bioRxivSearch

Biology subjects

Harris, R. S.

Publications and source records attributed to Harris, R. S..

5 recordsLinked to original sources

Gorilla APOBEC3 restricts SIVcpz and influences lentiviral evolution in great ape cross-species transmissions

Restriction factors including APOBEC3 family proteins have the potential prevent cross-species lentivirus transmissions. Such events as well as ensuing pathogenesis require the viral Vif protein to overcome/neutralize/degrade the APOBEC3 enzymes of the new host species. Previous investigations have focused on the molecular interaction between human APOBEC3s and HIV-1 Vif. However, the evolutionary interplay between lentiviruses and great ape (including human, chimpanzee and gorilla) APOBEC3s has not been fully investigated. Here we demonstrate that gorilla APOBEC3G plays a pivotal role in restricting lentiviral transmission from chimpanzee to gorilla. We also reveal that a sole amino acid substitution in Vif is sufficient to overcome the gorilla APOBEC3G-mediated species barrier. Moreover, the antiviral effects of gorilla APOBEC3D and APOBEC3F are considerably weaker than those of human and chimpanzee counterparts, which can result in the skewed evolution of great ape lentiviruses leading to HIV-1.\n\nHighlightsO_LISIVcpz requires M16E mutation in Vif to counteract gorilla A3G\nC_LIO_LIAcidic residue at position 16 of Vif is crucial to counteract gorilla A3G\nC_LIO_LIGorilla A3D and A3F poorly suppress lentiviral infectivity\nC_LIO_LISIVgor and related HIV-1s counteract human A3D and A3F independently of DRMR motif\nC_LI

microbiology

Non-B DNA affects polymerization speed and error rate in sequencers and living cells

DNA conformation may deviate from the classical B-form in ~13% of the human genome. Non-B DNA regulates many cellular processes; however, its effects on DNA polymerization speed and accuracy have not been investigated genome-wide. Such an inquiry is critical for understanding neurological diseases and cancer genome instability. Here we present the first simultaneous examination of DNA polymerization kinetics and errors in the human genome sequenced with Single-Molecule-Real-Time technology. We show that polymerization speed differs between non-B and B-DNA: it decelerates at G-quadruplexes and fluctuates periodically at disease-causing tandem repeats. Analyzing polymerization kinetics profiles, we predict and validate experimentally non-B DNA formation for a novel motif. We demonstrate that several non-B motifs affect sequencing errors (e.g., G-quadruplexes increase error rates) and that sequencing errors are positively associated with polymerase slowdown. Finally, we show that highly divergent G4 motifs have pronounced polymerization slowdown and high sequencing error rates, suggesting similar mechanisms for sequencing errors and germline mutations.

genomics

Increasing Cas9-mediated homology-directed repair efficiency through covalent tethering of DNA repair template

The CRISPR-Cas9 system is a powerful genome-editing tool in which a guide RNA targets Cas9 to a site in the genome where the Cas9 nuclease then induces a double stranded break (DSB)1,2. The potential of CRISPR-Cas9 to deliver precise genome editing is hindered by the low efficiency of homology-directed repair (HDR), which is required to incorporate a donor DNA template encoding desired genome edits near the DSB3,4. We present a strategy to enhance HDR efficiency by covalently tethering a single-stranded donor oligonucleotide (ssODN) to the Cas9/guide RNA ribonucleoprotein (RNP) complex via a fused HUH endonuclease5, thus spatially and temporally co-localizing the DSB machinery and donor DNA. We demonstrate up to an 8-fold enhancement of HDR using several editing assays, including repair of a frameshift and in-frame insertions of protein tags. The improved HDR efficiency is observed in multiple cell types and target loci, and is more pronounced at low RNP concentrations.

molecular biology

RecoverY : K-mer based read classification for Y-chromosome specific sequencing and assembly

MotivationThe haploid mammalian Y chromosome is usually under-represented in genome assemblies due to high repeat content and low depth due to its haploid nature. One strategy to ameliorate the low coverage of Y sequences is to experimentally enrich Y-specific material before assembly. Since the enrichment process is imperfect, algorithms are needed to identify putative Y-specific reads prior to downstream assembly. A strategy that uses k-mer abundances to identify such reads was used to assemble the gorilla Y (Tomaszkiewicz et al 2016). However, the strategy required the manual setting of key parameters, a time-consuming process leading to sub-optimal assemblies.\n\nResultsWe develop a method, RecoverY, that selects Y-specific reads by automatically choosing the abundance level at which a k-mer is deemed to originate from the Y. This algorithm uses prior knowledge about the Y chromosome of a related species or known Y transcript sequences. We evaluate RecoverY on both simulated and real data, for human and gorilla, and investigate its robustness to important parameters. We show that RecoverY leads to a vastly superior assembly compared to alternate strategies of filtering the reads or contigs. Compared to the preliminary strategy used in Tomaszkiewicz et al (2016), we achieve a 33% improvement in assembly size and a 20% improvement in the NG50, demonstrating the power of automatic parameter selection.\n\nAvailabilityOur tool RecoverY is freely available at https://github.com/makovalab-psu/RecoverY\n\nContactkmakova@bx.psu.edu, pashadag@cse.psu.edu\n\nSupplementary informationAttached as an additional file.

bioinformatics

AllSome Sequence Bloom Trees

The ubiquity of next generation sequencing has transformed the size and nature of many databases, pushing the boundaries of current indexing and searching methods. One particular example is a database of 2,652 human RNA-seq experiments uploaded to the Sequence Read Archive. Recently, Solomon and Kingsford proposed the Sequence Bloom Tree data structure and demonstrated how it can be used to accurately identify SRA samples that have a transcript of interest potentially expressed. In this paper, we propose an improvement called the AllSome Sequence Bloom Tree. Results show that our new data structure significantly improves performance, reducing the tree construction time by 52.7% and query time by 39 - 85%, with a price of up to 3x memory consumption during queries. Notably, it can query a batch of 198,074 queries in under 8 hours (compared to around two days previously) and a whole set of k-mers from a sequencing experiment (about 27 mil k-mers) in under 11 minutes.

bioinformatics