Search bioRxivSearch

Biology subjects

Greninger, A. L.

Publications and source records attributed to Greninger, A. L..

6 recordsLinked to original sources

VAPiD: a lightweight cross platform viral annotation pipeline and identification tool to facilitate virus genome submissions to NCBI GenBank

BackgroundWith sequencing technologies becoming cheaper and easier to use, more groups are able to obtain whole genome sequences of viruses of public health and scientific importance. Submission of genomic data to NCBI GenBank is a requirement prior to publication and plays a critical role in making scientific data publicly available.\n\nGenBank currently has automatic prokaryotic and eukaryotic genome annotation pipelines but has no viral annotation pipeline beyond influenza virus. Annotation and submission of viral genome sequence is a non-trivial task, especially for groups that do not routinely interact with GenBank for data submissions.\n\nResultsWe present Viral Annotation Pipeline and iDentification (VAPiD), a portable and lightweight command-line tool for annotation and GenBank deposition of viral genomes. VAPiD supports annotation of nearly all unsegmented viral genomes. The pipeline has been validated on human immunodeficiency virus, human parainfluenza virus 1-4, human metapneumovirus, human coronaviruses (229E/OC43/NL63/HKU1/SARS/MERS), human enteroviruses/rhinoviruses, measles virus, mumps virus, Hepatitis A-E Virus, Chikungunya virus, dengue virus, and West Nile virus, as well the human polyomaviruses BK/JC/MCV, human adenoviruses, and human papillomaviruses. The program can handle individual or batch submissions of different viruses to GenBank and correctly annotates multiple viruses, including those that contain ribosomal slippage or RNA editing without prior knowledge of the virus to be annotated. VAPiD is programmed in Python and is compatible with Windows, Linux, and Mac OS systems.\n\nConclusionsWe have created a portable, lightweight, user-friendly, internet-enabled, open-source, command-line genome annotation and submission package to facilitate virus genome submissions to NCBI GenBank. Instructions for downloading and installing VAPiD can be found at https://github.com/rcs333/VAPiD.

bioinformatics

Limited marginal utility of deep sequencing for HIV drug resistance testing in the age of integrase inhibitors

HIV drug resistance genotyping is a critical tool in the clinical management of HIV infections. Although resistance genotyping has traditionally been conducted using Sanger sequencing, next-generation sequencing (NGS) is emerging as a powerful tool due to its ability to detect lower frequency alleles. However, the value added from NGS approaches to antiviral resistance testing remains to be demonstrated. We compared the variant detection capacity of NGS versus Sanger sequencing methods for resistance genotyping of 144 drug resistance tests (105 protease-reverse transcriptase tests and 39 integrase tests) submitted to our clinical virology laboratory over a four-month period in 2016 for Sanger-based HIV drug resistance testing. NGS detected all true high frequency drug resistance mutations (>20% frequency) found by Sanger sequencing, with greater accuracy in one instance of a Sanger-detected false positive. Freely available online NGS variant callers Hydra and PASeq were superior to Sanger methods for interpretations of allele linkage and automated variant calling. NGS additionally detected low frequency mutations (1-20% frequency) associated with higher levels of drug resistance in 30/105 (29%) of protease-reverse transcriptase tests and 4/39 (10%) of integrase tests. Clinical follow-up of 69 individuals for a median of 674 days found no difference in rates of virological failure between individuals with and without low frequency mutations, although rates of virological failure were higher for individuals with drug-relevant low frequency mutations. However, all 27 individuals who experienced virological failure reported poor adherence to their drug regimen during preceding follow-up time, and all 19 who subsequently improved their adherence achieved viral suppression at later time points consistent with a lack of clinical resistance. In conclusion, in a population with low antiviral resistance emergence, NGS methods detected numerous instances of minor alleles that did not result in subsequent bona fide virological failure due to antiviral resistance.\n\nImportanceGenotypic antiviral resistance testing for HIV is an essential component of the clinical microbiology and virology laboratory. Next-generation sequencing (NGS) has emerged as a powerful tool for the detection of low frequency sequence variants (allele frequencies <20%). Whether detecting these low frequency mutations in HIV contributes to improved patient health, however, remains unclear. We compared NGS to conventional Sanger sequencing for detecting resistance mutations for 144 HIV drug resistance tests submitted to our clinical virology laboratory and detected low frequency mutations in 24% of tests. Over approximately two years of follow-up for 69 patients for which we had access to electronic health records, no patients had virological failure due to antiviral resistance. Instead, virological failure was entirely explained by medication non-adherence. While larger studies are required, we suggest that detection of low frequency variants by NGS presents limited marginal clinical utility when compared to standard of care.

epidemiology

Ultrasensitive capture of human herpes simplex virus genomes directly from clinical samples reveals extraordinarily limited evolution in cell culture

Herpes simplex viruses (HSV) are difficult to sequence due to their large DNA genome, high GC content, and the presence of repeats. To date, most HSV genomes have been recovered from culture isolates, raising concern that these genomes may not accurately represent circulating clinical strains. We report the development and validation of a DNA oligonucleotide hybridization panel to recover near complete HSV genomes at abundances up to 50,000-fold lower than previously reported. Using copy number information on herpesvirus and host DNA background via quantitative PCR, we developed a protocol for pooling for cost-effective recovery of more than 50 HSV-1 or HSV-2 genomes per MiSeq run. We demonstrate the ability to recover >99% of the HSV genome at >100X coverage in 72 hours at viral loads that allow whole genome recovery from latently-infected ganglia. We also report a new computational pipeline for rapid HSV genome assembly and annotation. Using the above tools and a series of 17 HSV-1-positive clinical swabs sent to our laboratory for viral isolation, we show limited evolution of HSV-1 during viral isolation in human fibroblast cells compared to the original clinical samples. Our data indicate that previous studies using low passage clinical isolates of herpes simplex viruses are reflective of the viral sequences present in the lesion and thus can be used in phylogenetic analyses. We also detect superinfection within a single sample with unrelated HSV-1 strains recovered from separate oral lesions in an immunosuppressed patient during a 2.5-week period, illustrating the power of direct-from-specimen sequencing of HSV.\n\nImportanceHerpes simplex viruses affect more than 4 billion people across the globe, constituting a large burden of disease. Understanding global diversity of herpes simplex viruses is important for diagnostics and therapeutics as well as cure research and tracking transmission among humans. To date, most HSV genomics has been performed on culture isolates and DNA swabs with high quantities of virus. We describe the development of wet-lab and computational tools that enable the accurate sequencing of near-complete genomes of HSV-1 and HSV-2 directly from clinical specimens at abundances >50,000-fold lower than previously sequenced and at significantly reduced cost. We use these tools to profile circulating HSV-1 strains in the community and illustrate limited changes to the viral genome during the viral isolation process. These techniques enable cost-effective, rapid sequencing of HSV-1 and HSV-2 genomes that will help enable improved detection, surveillance, and control of this human pathogen.

genomics

Cooperating H3N2 influenza virus variants are not detectable in primary clinical samples

The high mutation rates of RNA viruses lead to rapid genetic diversification, which can enable cooperative interactions between variants in a viral population. We previously described two distinct variants of H3N2 influenza virus that cooperate in cell culture. These variants differ by a single mutation, D151G, in the neuraminidase protein. The D151G mutation reaches a stable frequency of about 50% when virus is passaged in cell culture. However, it is unclear whether selection for the cooperative benefits of D151G is a cell-culture phenomenon, or whether the mutation is also sometimes present at appreciable frequency in virus populations sampled directly from infected humans. Prior work has not detected D151G in unpassaged clinical samples, but these studies have used methods like Sanger sequencing and pyrosequencing that are relatively insensitive to low-frequency variation. We identified nine samples of human H3N2 influenza collected between 2013 to 2015 in which Sanger sequencing had detected a high frequency of the D151G mutation following one to three passages in cell culture. We deep-sequenced the unpassaged clinical samples to identify low-frequency viral variants. The frequency of D151G did not exceed the frequency of library preparation and sequencing errors in any of the sequenced samples. We conclude that passage in cell culture is primarily responsible for the frequent observations of D151G in recent H3N2 influenza strains.\n\nIMPORTANCEViruses mutate rapidly, and recent studies of RNA viruses have shown that related viral variants can sometimes cooperate to improve each others growth. We previously described two variants of H3N2 influenza virus that cooperate in cell culture. The mutation responsible for cooperation is often observed when human samples of influenza virus are grown in the lab before sequencing, but it is unclear whether the mutation also exists in human infections or is exclusively the result of lab passage. We identified nine human isolates of influenza that had developed the cooperating mutation after being grown in the lab, and performed highly sensitive deep-sequencing of the unpassaged clinical samples to determine whether the mutation existed in the original human infections. We found no evidence of the cooperating mutation in the unpassaged samples, suggesting that the cooperation primarily arises in laboratory conditions.

microbiology

Divergent In vitro MIC Characteristics and underlying isogenic mutations in host-specialized Pseudomonas aeruginosa

Clinical isolates of Pseudomonas aeruginosa (Pa) from patients with cystic fibrosis (CF) are known to differ from those associated with infections of non-CF hosts in colony morphology, drug susceptibility patterns, and genomic hypermutability. Although Pa isolates from CF have long been recognized for their overall higher resistance rate calculated generally by reduced \"percent susceptible\", this study takes the approach to compare and contrast Etest MIC distributions between two distinct cohorts of clinical strains (n=224 from 56 CF patients and n=130 from 68 non-CF patients respectively) isolated in 2013. Logarithmic transformed MIC (logMIC) values of 11 antimicrobial agents were compared between the two groups. CF isolates tended to produce heterogeneous and widely dispersed MICs compared to non-CF isolates. By applying a test for equality of variances, we were able to confirm that the MICs generated from CF isolates against 9 out of the 11 agents were significantly more dispersed than those from non-CF (p<0.02-<0.001). Quantile-quantiles plots indicated little agreement between the two cohorts of isolates. Based on whole genome sequencing of 19 representative CF Pa isolates, divergent gain- or loss-of-function mutations in efflux and porin genes and their regulators between isogenic or intra-clonal associates were evident. Not one, not a few, but the net effect all adaptive mutational changes in the genomes of CF Pa, both shared and unshared between isogenic strains, are responsible for the divergent heteroresistance patterns. Moreover, the isogenic variations are suggestive of a bacterial syntrophic lifestyle when \"lockedo inside a host focal airway environment over prolonged periods.\n\nSignificance statementBacterial heteroresistance is associated with niche specialized organisms interacting with host species for prolonged period of time, medically characterized by \"chronic focal infections\". A prime example is found in Pseudomonas aeruginosa isogenic/non-homogeneous isolates from patient airways with cystic fibrosis. The development of pseudomonal polarizing MICs in vitro to many actively used antimicrobial agents among isogenic isolates and \"Eagle-type\" heteroresistance patterns are common and characteristic. Widespread isogenic gene lesions were evident for defects in drug transporters, DNA mismatch repair, and many other structural or cellular functions--a result of pseudomonal symbiotic response to host selection. Co-isolation of extremely susceptible and resistant isogenic Pa strains suggests intra-airway evolution of a multicellular syntrophic bacterial lifestyle, which has laboratory interpretation and clinical treatment implications.

microbiology

Validation and Implementation of CLIA-Compliant Whole Genome Sequencing (WGS) in Public Health Laboratory

BackgroundPublic health microbiology laboratories (PHL) are at the cusp of unprecedented improvements in pathogen identification, antibiotic resistance detection, and outbreak investigation by using whole genome sequencing (WGS). However, considerable challenges remain due to the lack of common standards.\n\nObjectives1) Establish the performance specifications of WGS applications used in PHL to conform with CLIA (Clinical Laboratory Improvements Act) guidelines for laboratory developed tests (LDT), 2) Develop quality assurance (QA) and quality control (QC) measures, 3) Establish reporting language for end users with or without WGS expertise, 4) Create a validation set of microorganisms to be used for future validations of WGS platforms and multi-laboratory comparisons and, 5) Create modular templates for the validation of different sequencing platforms.\n\nMethodsMiSeq Sequencer and Illumina chemistry (Illumina, Inc.) were used to generate genomes for 34 bacterial isolates with genome sizes from 1.8 to 4.7 Mb and wide range of GC content (32.1%-66.1%). A customized CLCbio Genomics Workbench - shell script bioinformatics pipeline was used for the data analysis.\n\nResultsWe developed a validation panel comprising ten Enterobacteriaceae isolates, five gram-positive cocci, five gram-negative non-fermenting species, nine Mycobacterium tuberculosis, and five miscellaneous bacteria; the set represented typical workflow in the PHL. The accuracy of MiSeq platform for individual base calling was >99.9% with similar results shown for reproducibility/repeatability of genome-wide base calling. The accuracy of phylogenetic analysis was 100%. The specificity and sensitivity inferred from MLST and genotyping tests were 100%. A test report format was developed for the end users with and without WGS knowledge.\n\nConclusionWGS was validated for routine use in PHL according to CLIA guidelines for LDTs. The validation panel, sequencing analytics, and raw sequences will be available for future multi-laboratory comparisons of WGS in PHL. Additionally, the WGS performance specifications and modular validation template are likely to be adaptable for the validation of other platforms and reagents kits.

microbiology