Search bioRxivSearch

Biology subjects

Nguyen, T.

Publications and source records attributed to Nguyen, T..

18 recordsLinked to original sources

Aging Boosts Antiviral CD8+T Cell Memory Through Improved Engagement Of Diversified Recall Response Determinants

The determinants of protective CD8+ memory T cell (CD8+TM) immunity remain incompletely defined and may in fact constitute an evolving agency as aging CD8+TM progressively acquire enhanced rather than impaired recall capacities. Here, we show that old as compared to young antiviral CD8+TM more effectively harness disparate molecular processes (cytokine signaling, trafficking, effector functions, and co-stimulation/inhibition) that in concert confer greater secondary reactivity. The relative reliance on these pathways is contingent on the nature of the secondary challenge (greater for chronic than acute viral infections) and over time, aging CD8+TM re-establish a dependence on the same accessory signals required for effective priming of naive CD8+T cells in the first place. Thus, our findings are consistent with the recently proposed \"rebound model\" that stipulates a gradual alignment of naive and CD8+TM properties, and identify a diversified collection of potential targets that may be exploited for the therapeutic modulation of CD8+TM immunity.

immunology

Cancer as a tissue anomaly: classifying tumor transcriptomes based only on healthy data

Since the turn of the century, researchers have sought to diagnose cancer based on gene expression signatures measured from the blood or biopsy as biomarkers. This task, known as classification, is typically solved using a suite of algorithms that learn a mathematical rule capable of discriminating one group (e.g., cases) from another (e.g., controls). However, discriminatory methods can only identify cancerous samples that resemble those that the algorithm already saw during training. As such, we argue that discriminatory methods are fundamentally ill-suited for the classification of cancer: because the possibility space of cancer is definitively large, the existence of a one-of-a-kind gene expression signature becomes very likely. Instead, we propose using an established surveillance method that detects anomalous samples based on their deviation from a learned normal steady-state structure. By transferring this method to transcriptomic data, we can create an anomaly detector for tissue transcriptomes, a \"tissue detector\", that is capable of identifying cancer without ever seeing a single cancer example. Using models trained on normal GTEx samples, we show that our \"tissue detector\" can accurately classify TCGA samples as normal or cancerous and that its performance is further improved by including more normal samples in the training set. We conclude this report by emphasizing the conceptual advantages of anomaly detection and by highlighting future directions for this field of study.

bioinformatics

Generation of human neural retina transcriptome atlas by single cell RNA sequencing

The retina is a highly specialized neural tissue that senses light and initiates image processing. Although the functional organisation of specific cells within the retina has been well-studied, the molecular profile of many cell types remains unclear in humans. To comprehensively profile cell types in the human retina, we performed single cell RNA-sequencing on 20,009 cells obtained post-mortem from three donors and compiled a reference transcriptome atlas. Using unsupervised clustering analysis, we identified 18 transcriptionally distinct clusters representing all known retinal cells: rod photoreceptors, cone photoreceptors, Muller glia cells, bipolar cells, amacrine cells, retinal ganglion cells, horizontal cells, retinal astrocytes and microglia. Notably, our data captured molecular profiles for healthy and early degenerating rod photoreceptors, and revealed a novel role of MALAT1 in putative rod degeneration. We also demonstrated the use of this retina transcriptome atlas to benchmark pluripotent stem cell-derived cone photoreceptors and an adult Muller glia cell line. This work provides an important reference with unprecedented insights into the transcriptional landscape of human retinal cells, which is fundamental to our understanding of retinal biology and disease.

systems biology

A simple cloning-free method to efficiently induce gene expression using CRISPR/Cas9

Gain-of-function studies often require the tedious cloning of transgene cDNA into vectors for overexpression beyond the physiological expression levels. The rapid development of CRISPR/Cas technology presents promising opportunities to address these issues. Here we report a simple, cloning-free method to induce gene expression at endogenous locus using CRISPR/Cas9 activators. Our strategy utilises synthesized sgRNA expression cassettes to direct a nuclease-null Cas9 complex fused with transcriptional activators (VP64, p65 and Rta) for site-specific induction of endogenous genes. This strategy allows rapid initiation of gain-of-function studies in the same day. Using this cloning-free approach, we tested two CRISPR activation systems, dSpCas9VPR and dSaCas9VPR, for induction of multiple genes in human and rat cells. Our results showed that both CRISPR activators allow efficient induction of six different neural development genes (CRX, RORB, RAX, OTX2, ASCL1 and NEUROD1) in human cells, whereas the rat cells exhibit a more variable and less efficient levels of gene induction, as observed in three different genes (Ascl1, Neurod1, Nrl). Altogether, this study provides a simple method to efficiently activate endogenous gene expression using CRISPR/Cas9 activators, which can be applies as a rapid workflow to initiate gain-of-function studies for a range of molecular and cell biology disciplines.

molecular biology

CEACAM1 regulates the IL-6 mediated fever response to LPS through the RP105 receptor in murine monocytes

Systemic inflammation and the fever response to pathogens are coordinately regulated by IL-6 and IL-1{beta}. We previously showed that CEACAM1 regulates the LPS driven expression of IL-1{beta} in murine neutrophils through its ITIM receptor. We now show that the prompt secretion of IL-6 in response to LPS is regulated by CEACAM1 expression on bone marrow monocytes. Ceacam1-/- mice over-produce IL-6 in response to an i.p. LPS challenge, resulting in prolonged surface temperature depression and overt diarrhea compared to their wild type counterparts. Intraperitoneal injection of a 64Cu-labeled LPS, PET imaging agent shows confined localization to the peritoneal cavity, and fluorescent labeled LPS is taken up by myeloid splenocytes and muscle endothelial cells. While bone marrow monocytes and their progenitors (CD11b+Ly6G-) express IL-6 in the early response (<2 hours) to LPS in vitro, these cells are not detected in the bone marrow after in vivo LPS treatment due to their rapid and complete mobilization to the periphery. Notably, tissue macrophages are not involved in the early IL-6 response to LPS. In contrast to human monocytes, TLR4 is not expressed on murine bone marrow monocytes. Instead, the alternative LPS receptor RP105 is expressed and recruits MD1, CD14, Src, VAV1 and {beta}-actin in response to LPS to produce IL-6. CEACAM1 negatively regulates RP105 signaling in monocytes by recruitment of SHP-1, resulting in the sequestration of pVAV1 and {beta}-actin from RP105. This novel pathway and regulation of IL-6 producing by CEACAM1 defines a novel role for monocytes in the fever of mice to LPS.\n\nAUTHOR SUMMARYFever is one of the most common signs of the immune response to pathogens. The fever response to LPS or endotoxin of gram-negative bacteria is mediated by the combined action of two cytokines, IL-1{beta} and IL-6. Regulation of their production in response to LPS is an important area of investigation. While we previously showed that the regulation of IL-1{beta} production in neutrophils is through the lymphocyte receptor CEACAM1, we were interested if a similar mechanism operated for IL-6. Using a mouse model in which the CEACAM1 gene was knocked out, we show that IL-6 is over-produced compared to normal mice, and that monocytes, rather than neutrophils were the principal IL-6 producing cells. Surprisingly, murine monocytes do not express TLR4, the most commonly studied receptor for LPS, but instead express the low affinity LPS receptor, RP105, a receptor common expressed on B-cells. Furthermore, we show that bone marrow monocytes are rapidly released into the blood and home to tissues throughout the body in response to LPS. These findings explain much of the confusion in the literature concerning the immediate source of IL-6 and the distinct differences between murine and human monocytes in their in responses to LPS.

immunology

Improving the classification of neuropsychiatric conditions using gene ontology terms as features

Although neuropsychiatric disorders have a well-established genetic background, their specific molecular foundations remain elusive. This has prompted many investigators to design studies that identify explanatory biomarkers, and then use these biomarkers to predict clinical outcomes. One approach involves using machine learning algorithms to classify patients based on blood mRNA expression from high-throughput transcriptomic assays. However, these endeavours typically fail to achieve the high level of performance, stability, and generalizability required for clinical translation. Moreover, these classifiers can lack interpretability because informative genes do not necessarily have relevance to researchers. For this study, we hypothesized that annotation-based classifiers can improve classification performance, stability, generalizability, and interpretability. To this end, we evaluated the performance of four classification algorithms on six neuropsychiatric data sets using four annotation databases. Our results suggest that the Gene Ontology Biological Process database can transform gene expression into an annotation-based feature space that improves the performance and stability of blood-based classifiers for neuropsychiatric conditions. We also show how annotation features can improve the interpretability of classifiers: since annotation databases are often used to assign biological importance to genes, annotation-based classifiers are easy to interpret because the biological importance of the features are the features themselves. We found that using annotations as features improves the performance and stability of classifiers. We also noted that the top ranked annotations tend contain the top ranked genes, suggesting that the most predictive annotations are a superset of the most predictive genes. Based on this, and the fact that annotations are used routinely to assign biological importance to genetic data, we recommend transforming gene-level expression into annotation-level expression prior to the classification of neuropsychiatric conditions.

bioinformatics

Inhibition of thrombocyte activation restores protective immunity to mycobacterial infection

Infection-induced thrombocytosis is a clinically important complication of tuberculosis (TB). Recent studies have separately highlighted a correlation of platelet activation with TB severity and utility of aspirin as a host-directed therapy for TB that modulates the inflammatory response. Here we investigate the possibility that the beneficial effects of aspirin are related to an anti-platelet mode of action. We utilize the zebrafish-Mycobacterium marinum model to show mycobacteria drive host hemostasis through the formation of granulomas. Treatment of infected zebrafish with aspirin or platelet-specific glycoprotein IIb/IIIa inhibitors reduced mycobacterial burden demonstrating a detrimental role for infection-induced thrombocyte activation. We found platelet inhibition reduced thrombocyte-macrophage interactions and restored indices of macrophage-mediated immunity to mycobacterial infection. Pathological thrombocyte activation and granuloma formation were found to be intrinsically linked illustrating a bidirectional relationship between host hemostasis and TB pathogenesis. Our study illuminates platelet activation as an efficacious target of anti-platelets drugs including aspirin, a widely available and affordable host-directed therapy candidate for tuberculosis.\n\nKey PointsO_LIInhibition of thrombocyte activation improves control of mycobacterial infection.\nC_LIO_LIInhibition of thrombocyte activation reduces thrombocyte-macrophage interactions and improves indices of macrophage immune function against mycobacterial infection.\nC_LI

immunology

Genetic regulatory mechanisms of smooth muscle cells map to coronary artery disease risk loci

Coronary artery disease (CAD) is the leading cause of death globally. Genome-wide association studies (GWAS) have identified more than 95 independent loci that influence CAD risk, most of which reside in non-coding regions of the genome. To interpret these loci, we generated transcriptome and whole-genome datasets using human coronary artery smooth muscle cells (HCASMC) from 52 unrelated donors, as well as epigenomic datasets using ATAC-seq on a subset of 8 donors. Through systematic comparison with publicly available datasets from GTEx and ENCODE projects, we identified transcriptomic, epigenetic, and genetic regulatory mechanisms specific to HCASMC. We assessed the relevance of HCASMC to CAD risk using transcriptomic and epigenomic level analyses. By jointly modeling eQTL and GWAS datasets, we identified five genes (SIPA1, TCF21, SMAD3, FES, and PDGFRA) that modulate CAD risk through HCASMC, all of which have relevant functional roles in vascular remodeling. Comparison with GTEx data suggests that SIPA1 and PDGFRA influence CAD risk predominantly through HCASMC, while other annotated genes may have multiple cell and tissue targets. Together, these results provide new tissue-specific and mechanistic insights into the regulation of a critical vascular cell type associated with CAD in human populations.

genomics

Solving for X: evidence for sex-specific autism biomarkers across multiple transcriptomic studies

Autism spectrum disorder (ASD) is a markedly heterogeneous condition with a varied phenotypic presentation. Its high concordance among siblings, as well as its clear association with specific genetic disorders, both point to a strong genetic etiology. However, the molecular basis of ASD is still poorly understood, although recent studies point to the existence of sex-specific ASD pathophysiologies and biomarkers. Despite this, little is known about how exactly sex influences the gene expression signatures of ASD probands. In an effort to identify sex-dependent biomarkers (and characterise their function), we present an analysis of a single paired-end post-mortem brain RNA-Seq data set and a meta-analysis of six blood-based microarray data sets. Here, we identify several genes with sex-dependent dysregulation, and many more with sex-independent dysregulation. Moreover, through pathway analysis, we find that these sex-independent biomarkers have substantially different biological roles than the sex-dependent biomarkers, and that some of these pathways are ubiquitously dysregulated in both post-mortem brain and blood. We conclude by synthesizing the discovered biomarker profiles with the extant literature, by highlighting the advantage of studying sex-specific dysregulation directly, and by making a call for new transcriptomic data that comprise large female cohorts.

neuroscience

Lost in translation: egg transcriptome reveals molecular signature to predict developmental success and novel maternal-effect genes

BackgroundGood quality or developmentally competent eggs result in high survival of progeny. Previous research has shed light on factors that determine egg quality, however, large gaps remain. Initial development of the embryo relies on maternally-inherited molecules, such as transcripts, deposited in the egg, thus, they would likely reflect egg quality. We performed transcriptome analysis on zebrafish fertilized eggs of different quality from unrelated, wildtype couples to obtain a global portrait of the egg transcriptome to determine its association with developmental competence and to identify new candidate maternal-effect genes.\n\nResultsFifteen of the most differentially expressed genes (DEGs) were validated by quantitative real-time PCR. Gene ontology analysis showed that enriched terms included ribosomes and translation. In addition, statistical modeling using partial least squares regression and genetics algorithm also demonstrated that gene signatures from the transcriptomic data can be used to predict reproductive success. Among the validated DEGs, otulina and slc29a1a were found to be increased in good quality eggs and to be predominantly localized in the ovaries. CRISPR/Cas9 knockout mutants of each gene revealed remarkable subfertility whereby the majority of their embryos were unfertilizable. The Wnt pathway appeared to be dysregulated in the otulina knockout-derived eggs.\n\nConclusionsOur novel findings suggested that even in varying quality of eggs due to heterogeneous causes from unrelated wildtype couples, gene signatures exist in the egg transcriptome, which can be used to predict developmental competence. Further, transcriptomic profiling revealed two new potential maternal-effect genes that have essential roles in vertebrate reproduction.

genomics

WSL5, a pentatricopeptide repeat protein, is essential for chloroplast biogenesis in rice under cold stress

AbstactChloroplasts play an essential role in plant growth and development, and cold has a great effect on chloroplast development. Although many genes or regulators involved in chloroplast biogenesis and development have been isolated and characterized, identification of novel components associated with cold is still lacking. In this study, we reported the functional characterization of white stripe leaf 5 (wsl5) mutant in rice. The mutant developed white-striped leaves during early leaf development and was albinic when planted under cold stress. Genetic and molecular analysis revealed that WSL5 encodes a novel chloroplast-targeted pentatricopeptide repeat protein. RNA-seq analysis showed that expression of nuclear-encoded photosynthetic genes in the mutant was significantly repressed, and expression of many chloroplast-encoded genes was also significantly changed. Notably, the WSL5 mutation caused defects in editing of rpl2 and atpA, and in splicing of rpl2 and rps12. Chloroplast ribosome biogenesis was impaired under cold stress. We propose that WSL5 is required for normal chloroplast development in rice under cold stress.

genetics

Community assessment of cancer drug combination screens identifies strategies for synergy prediction

The effectiveness of most cancer targeted therapies is short lived since tumors evolve and develop resistance. Combinations of drugs offer the potential to overcome resistance, however the number of possible combinations is vast necessitating data-driven approaches to find optimal treatments tailored to a patients tumor. AstraZeneca carried out 11,576 experiments on 910 drug combinations across 85 cancer cell lines, recapitulating in vivo response profiles. These data, the largest openly available screen, were hosted by DREAM alongside deep molecular characterization from the Sanger Institute for a Challenge to computationally predict synergistic drug pairs and associated biomarkers. 160 teams participated to provide the most comprehensive methodological development and subsequent benchmarking to date. Winning methods incorporated prior knowledge of putative drug target interactions. For >60% of drug combinations synergy was reproducibly predicted with an accuracy matching biological replicate experiments, however 20% of drug combinations were poorly predicted by all methods. Genomic rationale for synergy predictions were identified, including antagonism unique to combined PIK3CB/D inhibition with the ADAM17 inhibitor where synergy is seen with other PI3K pathway inhibitors. All data, methods and code are freely available as a resource to the community.

bioinformatics

IRF4 haploinsufficiency in a family with Whipples disease

The pathogenesis of Whipples disease (WD) remains largely unknown, as WD strikes only a very small minority of the individuals infected with Tropheryma whipplei (Tw). Asymptomatic carriage of Tw is less rare. We studied a large multiplex French kindred, containing four otherwise healthy WD patients (mean age: 76.7 years) and five healthy carriers of Tw (mean age: 55 years). We used a strategy combining genome-wide linkage analysis and whole-exome sequencing to test the hypothesis that WD is inherited in an autosomal dominant (AD) manner, with age-dependent incomplete penetrance. WD was linked to 12 genomic regions covering 27 megabases in the four patients. These regions contained only one very rare non-synonymous variation: the R98W variant of IRF4. The five Tw carriers were heterozygous for R98W. Interferon regulatory factor 4 (IRF4) is a transcription factor with pleiotropic roles in immunity. We showed that R98W was a loss-of-function allele, like only five other exceedingly rare IRF4 alleles of a total of 39 rare and common non-synonymous alleles tested. Furthermore, heterozygosity for R98W led to a distinctive pattern of transcription in leukocytes following stimulation with BCG or Tw. Finally, we found that IRF4 had evolved under purifying selection and that R98W was not dominant-negative, suggesting that the IRF4 deficiency in this kindred was due to haploinsufficiency. Overall, haploinsufficiency at the IRF4 locus selectively underlies WD in this multiplex kindred. This deficiency displays AD inheritance with incomplete penetrance, and chronic carriage probably precedes WD by several decades in Tw-infected heterozygotes.

immunology

Aging Of Antiviral CD8+ Memory T Cells Fosters Increased Survival, Metabolic Adaptations And Lymphoid Tissue Homing

Aging of established antiviral T cell memory fosters a series of progressive adaptations that paradoxically improve rather than compromise protective CD8+T cell immunity. We now provide evidence that this gradual evolution, the pace of which is contingent on the precise context of the primary response, also impinges on the molecular mechanisms that regulate CD8+ memory T cell (CD8+TM) homeostasis. Over time, CD8+TM become more resistant to apoptosis and acquire enhanced cytokine responsiveness without adjusting their homeostatic proliferation rates; concurrent metabolic adaptations promote increased CD8+TM quiescence and fitness but also impart the re-acquisition of a partial effector-like metabolic profile; and a gradual redistribution of aging CD8+TM from blood and nonlymphoid tissues to lymphatic organs results in CD8+TM accumulations in bone marrow, splenic white pulp and particularly lymph nodes. Altogether, these data demonstrate how temporal alterations of fundamental homeostatic determinants converge to render aged CD8+TM poised for greater recall responses.\n\nABBREVIATIONST cell subsets

immunology

Genomic and proteomic analysis of Human herpesvirus 6 reveals distinct clustering of acute versus inherited forms and reannotation of reference strain

Human herpesvirus-6A and -6B (HHV-6) are betaherpesviruses that reach >90% seroprevalence in the adult population. Unique among human herpesviruses, HHV-6 can integrate into the subtelomeric regions of human chromosomes; when this occurs in germ line cells it causes a condition called inherited chromosomally integrated HHV-6 (iciHHV-6). To date, only two complete genomes are available for HHV-6B. Using a custom capture panel for HHV-6B, we report near-complete genomes from 61 isolates of HHV-6B from active infections (20 from Japan, 35 from New York state, and 6 from Uganda), and 64 strains of iciHHV-6B (mostly from North America). We also report partial genome sequences from 10 strains of iciHHV-6A. Although the overall sequence diversity of HHV-6 is limited relative to other human herpesviruses, our sequencing identified geographical clustering of HHV-6B sequences from active infections, as well as evidence of recombination among HHV-6B strains. One strain of active HHV-6B was more divergent than any other HHV-6B previously sequenced. In contrast to the active infections, sequences from iciHHV-6 cases showed reduced sequence diversity. Strikingly, multiple iciHHV-6B sequences from unrelated individuals were found to be completely identical, consistent with a founder effect. However, several iciHHV-6B strains intermingled with strains from active pediatric infection, consistent with the hypothesis that intermittent de novo integration into host germline cells can occur during active infection Comparative genomic analysis of the newly sequenced strains revealed numerous instances where conflicting annotations between the two existing reference genomes could be resolved. Combining these findings with transcriptome sequencing and shotgun proteomics, we reannotated the HHV-6B genome and found multiple instances of novel splicing and genes that hitherto had gone unannotated. The results presented here constitute a significant genomic resource for future studies on the detection, diversity, and control of HHV-6.\n\nAuthor SummaryHHV-6 is a ubiquitous large DNA virus that is the most common cause of febrile seizures and reactivates in allogeneic stem cell patients. It also has the unique ability among human herpesviruses to be integrated into the genome of every cell via integration in the germ line, a condition called inherited chromosomally integrated (ici)HHV-6, which affects approximately 1% of the population. To date, very little is known about the comparative genomics of HHV-6. We sequenced 61 isolates of HHV-6B from active infections, 64 strains of iciHHV-6B, and 10 strains of iciHHV-6A. We found geographic clustering of HHV-6B strains from active infections. In contrast, iciHHV-6B had reduced sequence diversity, with many identical sequences of iciHHV-6 found in individuals not known to share recent common ancestry, consistent with a founder effect from a remote common ancestor with iciHHV-6. We also combined our genomic analysis with transcriptome sequencing and shotgun proteomics to correct previous misannotations of the HHV-6 genome.

microbiology

Enhancer connectome in primary human cells reveals target genes of disease-associated DNA elements

The challenge of linking intergenic mutations to target genes has limited molecular understanding of diverse human diseases. Here, we show H3K27ac HiChIP generates high-resolution contact maps of active enhancers and target genes in rare primary human T cell subtypes and coronary artery smooth muscle cells. Differentiation of naive T cells to either T helper 17 cells or regulatory T cells create subtype-specific enhancer-promoter interactions, specifically at regions of shared DNA accessibility. These data provide a principled means of assigning molecular functions to autoimmune and cardiovascular disease risk variants, linking hundreds of noncoding variants to putative gene targets. Target genes identified with HiChIP are further supported by CRISPR interference and activation at linked enhancers, by the presence of expression quantitative trait loci, and by allele-specific enhancer loops in patient-derived primary cells. The majority of disease-associated enhancers contact genes beyond the nearest gene in the linear genome, leading to a four-fold increase of potential target genes for autoimmune and cardiovascular diseases.

genomics

Evolution of gene expression after whole-genome duplication: new insights from the spotted gar genome

Whole genome duplications (WGD) are important evolutionary events. Our understanding of underlying mechanisms, including the evolution of duplicated genes after WGD, however remains incomplete. Teleost fish experienced a common WGD (teleost-specific genome duplication, or TGD) followed by a dramatic adaptive radiation leading to more than half of all vertebrate species. The analysis of gene expression patterns following TGD at the genome level has been limited by the lack of suitable genomic resources. The recent concomitant release of the genome sequence of spotted gar (a representative of holosteans, the closest lineage of teleosts that lacks the TGD) and the tissue-specific gene expression repertoires of over 20 holostean and teleostean fish species, including spotted gar, zebrafish and medaka (the PhyloFish project), offered a unique opportunity to study the evolution of gene expression following TGD in teleosts. We show that most TGD duplicates gained their current status (loss of one duplicate gene or retention of both duplicates) relatively rapidly after TGD (i.e. prior to the divergence of medaka and zebrafish lineages). The loss of one duplicate is the most common fate after TGD with a probability of approximately 80%. In addition, the fate of duplicate genes after TGD, including subfunctionalization, neofunctionalization, or retention of two similar copies occurred not only before, but also after the radiation of species tested, in consistency with a role of the TGD in speciation and/or evolution of gene function. Finally, we report novel cases of TGD ohnolog subfunctionalization and neofunctionalization that further illustrate the importance of these processes.

evolutionary biology

Genome-Wide Comparison Of Toxigenic And Non-Toxigenic Corynebacterium diphtheriae Isolates Identifies Differences In The Pan Genomes Between Respiratory And Cutaneous Strains

ObjectivesCorynebacterium diphtheriae is the main etiological agent of diphtheria, a global disease causing life-threatening infections, particularly in infants and children. Vaccination with diphtheria toxoid protects against infection with potent toxin producing strains. However a growing number of apparently non-toxigenic but potentially invasive C. diphtheriae strains are identified in countries with low prevalence of diphtheria, raising key questions about genomic structures and population dynamics of the species.\n\nMethodsThis study examined genomic diversity among 47 C. diphtheriae isolates collected in Australia over a 10-year period using whole genome sequencing. Phylogeny was determined using SNP-based mapping and genome wide analysis.\n\nResultsC. diphtheriae sequence type (ST) 32, a non-toxigenic ST with evidence of enhanced virulence that is also circulating in Europe, appears to be endemic in Australia. Isolates from temporospatially related patients displayed the same ST and similarity in their core genomes. The genome-wide analysis highlighted a role of pilins, adhesion factors and iron utilization in infections caused by toxigenic as well as non-toxigenic strains.\n\nConclusionsThe genomic diversity of toxigenic and non-toxigenic strains of C. diphtheriae in Australia suggests multiple local and overseas sources of infection and colonisation. Our findings suggest that regular genomic surveillance of co-circulating toxigenic and non-toxigenic C. diphtheriae can deliver highly nuanced data in order to inform targeted public health actions and policy for predicting the future impact of this highly successful pathogen.

microbiology