Search bioRxivSearch

Biology subjects

Albert, J.

Publications and source records attributed to Albert, J..

4 recordsLinked to original sources

Comprehensive analysis of intra- and interpatient evolution of enterovirus D68 by whole-genome deep sequencing

Worldwide outbreaks of enterovirus D68 (EV-D68) in 2014 and 2016 have caused widespread serious respiratory and neurological disease. To investigate diversity, spread, and evolution of EV-D68 we performed near full-length deep sequencing in 54 samples obtained in Sweden during the 2014 and 2016 outbreaks. In most samples, intrapatient variability was low and dominated by rare synonymous variants, but three patients showed evidence of dual infections with distinct EV-D68 variants from the same subclade. Interpatient evolution showed a very strong temporal signal, with an evolutionary rate of 0.0038{+/-}0.0001 substitutions per site and year. Phylogenetic trees reconstructed from the sequences suggest that EV-D68 was introduced into Stockholm several times during the 2016 outbreak and then spread locally. Putative neutralization targets in the BC- and DE-loops of the VP1 protein are slightly more diverse within-host and tend to undergo more frequent substitution than other genomic regions. However, evolution in these loops does not appear to have driven the emergence of the 2016 B3-subclade directly from the 2014 B1-subclade. Instead, the most recent ancestor of both clades was dated to 2009. The study provides a comprehensive description of the intra- and interpatient evolution of EV-D68, including the first report of intrapatient diversity and dual infections. The new data along with publicly available EV-D68 sequences are included in an interactive phylodynamic analysis on nextstrain.org/enterovirus/d68 to facilitate timely EV-D68 tracking in the future.xs

microbiology

Getting more from heterogeneous HIV-1 surveillance data in a high immigration country: estimation of incidence and undiagnosed population size using multiple biomarkers

BackgroundMost HIV infections originate from individuals who are undiagnosed and unaware of their infection. Estimation of this quantity from surveillance data is hard because there is incomplete knowledge about i) the time between infection and diagnosis (TI) for the general population and ii) the time between immigration and diagnosis for foreign-born persons.\n\nDevelopmentWe developed a new statistical method for estimating the number of undiagnosed people living with HIV (PLHIV) and the incidence of HIV-1 based on dynamic modeling of heterogenous HIV-1 surveillance data. We formulated a Bayesian non-linear mixed effects model using multiple biomarkers to estimate TI accounting for biomarker correlation and individual heterogeneities. We explicitly model the probability that an HIV-1 infected foreign-born person was infected either before or after immigration to distinguish between endogenous and exogeneous incidence. The incidence estimator allows for direct calculation of the number of undiagnosed persons.\n\nApplicationThe model was applied to surveillance data in Sweden. The dynamic biomarker model was trained on longitudinal data from 31 treatment-naive patients with well-defined TI, using CD4 counts, BED serology, polymorphisms in HIV-1 pol sequences, and testing history. The multiple-biomarker model was more accurate than single biomarkers (mean absolute error 1.01 vs [≥] 1.95). We estimate that 813 (95% CI 780-862) PLHIV were undiagnosed in 2015, representing a proportion of 10.8% (95% CI 10.4-11.3%) of all PLHIV.\n\nConclusionsThe proposed methodology will enhance the utility of standard surveillance data streams and will be useful to monitor progress towards and compliance with the 90-90-90 UNAIDS target.\n\nKey messagesO_LICombined heterogeneous HIV-1 surveillance data and biomarker data can be used to estimate both local incidence and the number of undiagnosed people living with HIV.\nC_LIO_LIExplicit modeling of the dynamics, heterogeneity, and correlation of multiple biomarkers over time improved estimation of time between infection and diagnosis.\nC_LIO_LIExplicit modeling of the probability that foreign-born persons were infected before or after immigration improves accuracy of estimates of endogenous incidence and undiagnosed persons living with HIV.\nC_LIO_LIThe endogenous incidence of HIV-1 in Sweden is declining, despite continued immigration of HIV-1 infected persons.\nC_LIO_LIThe proportion of undiagnosed PLHIV decreased over 2010-2015 and was estimated to be 10.8% (95% CI, 10.4-11.3%) in 2015.\nC_LI

epidemiology

Estimating Time Of HIV-1 Infection From Next-Generation Sequence Diversity

Estimating the time since infection (TI) in newly diagnosed HIV-1 patients is challenging, but important to understand the epidemiology of the infection. Here we explore the utility of virus diversity estimated by next-generation sequencing (NGS) as novel biomarker by using a recent genome-wide longitudinal dataset obtained from 11 untreated HIV-1-infected patients with known dates of infection. The results were validated on a second dataset from 31 patients.\n\nVirus diversity increased linearly with time, particularly at 3rd codon positions, with little inter-patient variation. The precision of the TI estimate improved with increasing sequencing depth, showing that diversity in NGS data yields superior estimates to the number of ambiguous sites in Sanger sequences, which is one of the alternative biomarkers. The full advantage of deep NGS was utilized with continuous diversity measures such as average pairwise distance or site entropy, rather than the fraction of polymorphic sites. The precision depended on the genomic region and codon position and was highest when 3rd codon positions in the entire pol gene were used. For these data, TI estimates had a mean absolute error of around 1 year. The error increased only slightly from around 0.6 years at a TI of 6 months to around 1.1 years at 6 years.\n\nOur results show that virus diversity determined by NGS can be used to estimate time since HIV-1 infection many years after the infection, in contrast to most alternative biomarkers. We provide the regression coefficients as well as web tool for TI estimation.\n\nAuthor summaryHIV-1 establishes a chronic infection, which may last for many years before the infected person is diagnosed. The resulting uncertainty in the date of infection leads to difficulties in estimating the number of infected but undiagnosed persons as well as the number of new infections, which is necessary for developing appropriate public health policies and interventions. Such estimates would be much easier if the time since HIV-1 infection for newly diagnosed cases could be accurately estimated. Three types of biomarkers have been shown to contain information about the time since HIV-1 infection, but unfortunately, they only distinguish between recent and long-term infections (concentration of HIV-1-specific antibodies) or are imprecise (immune status as measured by levels of CD4+ T-lymphocytes and viral sequence diversity measured by polymorphisms in Sanger sequences).\n\nIn this paper, we show that recent advances in sequencing technologies, i.e. the development of next generation sequencing, enable significantly more precise determination of the time since HIV-1 infection, even many years after the infection event. This is a significant advance which could translate into more effective HIV-1 prevention.

epidemiology

Easy and Accurate Reconstruction of Whole HIV Genomes from Short-Read Sequence Data

Next-generation sequencing has yet to be widely adopted for HIV. The difficulty of accurately reconstructing the consensus sequence of a quasispecies from reads (short fragments of DNA) in the presence of rapid between- and within-host evolution may have presented a barrier. In particular, mapping (aligning) reads to a reference sequence leads to biased loss of information; this bias can distort epidemiological and evolutionary conclusions. De novo assembly avoids this bias by effectively aligning the reads to themselves, producing a set of sequences called contigs. However contigs provide only a partial summary of the reads, misassembly may result in their having an incorrect structure, and no information is available at parts of the genome where contigs could not be assembled. To address these problems we developed the tool shiver to preprocess reads for quality and contamination, then map them to a reference tailored to the sample using corrected contigs supplemented with existing reference sequences. Run with two commands per sample, it can easily be used for large heterogeneous data sets. We use shiver to reconstruct the consensus sequence and minority variant information from paired-end short-read data produced with the Illumina platform, for 65 existing publicly available samples and 50 new samples. We show the systematic superiority of mapping to shivers constructed reference over mapping the same reads to the standard reference HXB2: an average of 29 bases per sample are called differently, of which 98.5% are supported by higher coverage. We also provide a practical guide to working with imperfect contigs.

bioinformatics