Search bioRxivSearch

Biology subjects

Kellam, P.

Publications and source records attributed to Kellam, P..

4 recordsLinked to original sources

Determining the Mutation Bias of Favipiravir in Influenza Using Next-generation Sequencing

AbstractFavipiravir is a broad spectrum antiviral drug that may be used to treat influenza. Previous research has identified that favipiravir likely acts as a mutagen but the precise mutation bias that favipiravir induces in influenza virus RNAs has not been described. Here, we use next-generation sequencing (NGS) with barcoding of individual RNA molecules to accurately and quantitatively detect favipiravir-induced mutations and to sample orders of magnitude more mutations than would be possible through Sanger sequencing. We demonstrate that favipiravir causes mutations and show that favipiravir primarily acts as a guanine analogue and secondarily as an adenine analogue resulting in the accumulation of transition mutations. We also use a standard NGS pipeline to show that the mutagenic effect of favipiravir can be measured by whole genome sequencing of virus.\n\nImportanceNew antiviral drugs are needed as a first line of defence in the event of a novel influenza pandemic. Favipiravir is a broad-spectrum antiviral which is effective against influenza. The exact mechanism of how favipiravir works to inhibit influenza is still unclear. We used next-generation sequencing (NGS) to demonstrate that favipiravir causes mutations in influenza RNA. The greater depth of NGS sequence information over traditional sequencing methods allowed us to precisely determine the bias of particular mutations caused by favipiravir. NGS can also be used in a standard diagnostic pipeline to show that favipiravir is acting on the virus by revealing the mutation bias pattern typical to the drug. Our work will aid in testing whether viruses are resistant to favipiravir and may help demonstrate the effect of favipiravir on viruses in a clinical setting. This will be important if favipiravir is used during a future influenza pandemic.

microbiology

Whole genome analysis of local Kenyan and global sequences unravels the epidemiological and molecular evolutionary dynamics of RSV genotype ON1 strains

The respiratory syncytial virus (RSV) group A variant with the 72-nucleotide duplication in the G gene, genotype ON1, was first detected in Kilifi in 2012 and has almost completely replaced previously circulating genotype GA2 strains. This replacement suggests some fitness advantage of ON1 over the GA2 viruses, and might be accompanied by important genomic substitutions in ON1 viruses. Close observation of such a new virus introduction over time provides an opportunity to better understand the transmission and evolutionary dynamics of the pathogen. We have generated and analyzed 184 RSV-A whole genome sequences (WGS) from Kilifi (Kenya) collected between 2011 and 2016, the first ON1 genomes from Africa and the largest collection globally from a single location. Phylogenetic analysis indicates that RSV-A transmission into this coastal Kenya location is characterized by multiple introductions of viral lineages from diverse origins but with varied success in local transmission. We identify signature amino acid substitutions between ON1 and GA2 viruses within genes encoding the surface proteins (G, F), polymerase (L) and matrix M2-1 proteins, some of which were identified as positively selected, and thereby provide an enhanced picture of RSV-A diversity. Furthermore, five of the eleven RSV open reading frames (ORF) (i.e. G, F, L, N and P), analyzed separately, formed distinct phylogenetic clusters for the two genotypes. This might suggest that coding regions outside of the most frequently studied G ORF play a role in the adaptation of RSV to host populations with the alternative possibility that some of the substitutions are nothing more than genetic hitchhikers. Our analysis provides insight into the epidemiological processes that define RSV spread, highlights the genetic substitutions that characterize emerging strains, and demonstrates the utility of large-scale WGS in molecular epidemiological studies.\n\nAuthor summaryRespiratory syncytial virus (RSV) is the leading viral cause of severe pneumonia and bronchiolitis among infants and children globally. No vaccine exists to date. The high genetic variability of this RNA virus, characterized by group (A or B), genotype (within group) and variant (within genotype) replacement in populations, may pose a challenge to effective vaccine design by enabling immune response escape. To date most sequence data exists for the highly variable G gene encoding the RSV attachment protein, and there is little globally-sampled RSV genomic data to provide a fine resolution of the epidemiology and evolutionary dynamics of the pathogen. Here we use long-term RSV surveillance in coastal Kenya to track the introduction, spread and evolution of a new RSV genotype known as ON1 (having a 72-nucleotide duplication in the G gene). We present a set of 184 RSV-A whole genomes, including 176 of RSV ON1 (the first from Africa), describe patterns of local ON1 spread and show genome-wide changes between the two major RSV-A genotypes that may define the pathogens adaptation to the host. These findings have implications for vaccine design and improved understanding of RSV epidemiology and evolution.

epidemiology

A High HIV-1 Strain Variability in London, UK, Revealed by Full-Genome Analysis: Results from the ICONIC Project

Background & MethodsThe ICONIC project has developed an automated high-throughput pipeline to generate HIV nearly full-length genomes (NFLG, i.e. from gag to nef) from next-generation sequencing (NGS) data. The pipeline was applied to 420 HIV samples collected at University College London Hospital and Barts Health NHS Trust (London) and sequenced using an Illumina MiSeq at the Wellcome Trust Sanger Institute (Cambridge). Consensus genomes were generated and subtyped using COMET, and unique recombinants were studied with jpHMM and SimPlot. Maximum-likelihood phylogenetic trees were constructed using RAxML to identify transmission networks using the Cluster Picker.\n\nResultsThe pipeline generated sequences of at least 1Kb of length (median=7.4Kb) for 375 out of the 420 samples (89%), with 174 (46.4%) being NFLG. A total of 365 sequences (169 of them NFLG) corresponded to unique subjects and were included in the down-stream analyses. The most frequent HIV subtypes were B (n=149, 40.8%) and C (n=77, 21.1%) and the circulating recombinant form CRF02_AG (n=32, 8.8%). We found 14 different CRFs (n=66, 18.1%) and multiple URFs (n=32, 8.8%) that involved recombination between 12 different subtypes/CRFs. The most frequent URFs were B/CRF01_AE (4 cases) and A1/D, B/C, and B/CRF02_AG (3 cases each). Most URFs (19/26, 73%) lacked breakpoints in the PR+RT pol region, rendering them undetectable if only that was sequenced. Twelve (37.5%) of the URFs could have emerged within the UK, whereas the rest were probably imported from sub-Saharan Africa, South East Asia and South America. For 2 URFs we found highly similar pol sequences circulating in the UK. We detected 31 phylogenetic clusters using the full dataset: 25 pairs (mostly subtypes B and C), 4 triplets and 2 quadruplets. Some of these were not consistent across different genes due to inter- and intra-subtype recombination. Clusters involved 70 sequences, 19.2% of the dataset.\n\nConclusionsThe initial analysis of genome sequences detected substantial hidden variability in the London HIV epidemic. Analysing full genome sequences, as opposed to only PR+RT, identified previously undetected recombinants. It provided a more reliable description of CRFs (that would be otherwise misclassified) and transmission clusters.

genomics

Easy and Accurate Reconstruction of Whole HIV Genomes from Short-Read Sequence Data

Next-generation sequencing has yet to be widely adopted for HIV. The difficulty of accurately reconstructing the consensus sequence of a quasispecies from reads (short fragments of DNA) in the presence of rapid between- and within-host evolution may have presented a barrier. In particular, mapping (aligning) reads to a reference sequence leads to biased loss of information; this bias can distort epidemiological and evolutionary conclusions. De novo assembly avoids this bias by effectively aligning the reads to themselves, producing a set of sequences called contigs. However contigs provide only a partial summary of the reads, misassembly may result in their having an incorrect structure, and no information is available at parts of the genome where contigs could not be assembled. To address these problems we developed the tool shiver to preprocess reads for quality and contamination, then map them to a reference tailored to the sample using corrected contigs supplemented with existing reference sequences. Run with two commands per sample, it can easily be used for large heterogeneous data sets. We use shiver to reconstruct the consensus sequence and minority variant information from paired-end short-read data produced with the Illumina platform, for 65 existing publicly available samples and 50 new samples. We show the systematic superiority of mapping to shivers constructed reference over mapping the same reads to the standard reference HXB2: an average of 29 bases per sample are called differently, of which 98.5% are supported by higher coverage. We also provide a practical guide to working with imperfect contigs.

bioinformatics