Search bioRxivSearch

Biology subjects

Gardy, J. L.

Publications and source records attributed to Gardy, J. L..

5 recordsLinked to original sources

The Integrated Rapid Infectious Disease Analysis (IRIDA) Platform

Whole genome sequencing (WGS) is a powerful tool for public health infectious disease investigations owing to its higher resolution, greater efficiency, and cost-effectiveness over traditional genotyping methods. Implementation of WGS in routine public health microbiology laboratories is impeded by a lack of user-friendly automated and semi-automated pipelines, restrictive jurisdictional data sharing policies, and the proliferation of non-interoperable analytical and reporting systems. To address these issues, we developed the Integrated Rapid Infectious Disease Analysis (IRIDA) platform (irida.ca), a user-friendly, decentralized, open-source bioinformatics and analytical web platform to support real-time infectious disease outbreak investigations using WGS data. Instances can be independently installed on local high-performance computing infrastructure, enabling private and secure data management and analyses according to organizational policies and governance. IRIDAs data management capabilities enable secure upload, storage and sharing of all WGS data and metadata. The core platform currently includes pipelines for quality control, assembly, annotation, variant detection, phylogenetic analysis, in silico serotyping, multi-locus sequence typing, and genome distance calculation. Analysis pipeline results can be visualized within the platform through dynamic line lists and integrated phylogenomic clustering for research and discovery, and for enhancing decision-making support and hypothesis generation in epidemiological investigations. Communication and data exchange between instances are provided through customizable access controls. IRIDA complements centralized systems, empowering local analytics and visualizations for genomics-based microbial pathogen investigations. IRIDA is currently transforming the Canadian public health ecosystem and is freely available at https://github.com/phac-nml/irida and www.irida.ca.\n\nImpact StatementWhole genome sequencing (WGS) is revolutionizing infectious disease analysis and surveillance due to its cost effectiveness, utility, and improved analytical power. To date, no \"one-size-fits-all\" genomics platform has been universally adopted, owing to differences in national (and regional) health information systems, data sharing policies, computational infrastructures, lack of interoperability and prohibitive costs. The Integrated Rapid Infectious Disease Analysis (IRIDA) platform is a user-friendly, decentralized, open-source bioinformatics and analytical web platform developed to support real-time infectious disease outbreak investigations using WGS data. IRIDA empowers public health, regulatory and clinical microbiology laboratory personnel to better incorporate WGS technology into routine operations by shielding them from the computational and analytical complexities of big data genomics. IRIDA is now routinely used as part of a validated suite of tools to support outbreak investigations in Canada. While IRIDA was designed to serve the needs of the Canadian public health system, it is generally applicable to any public health and multi-jurisdictional environment. IRIDA enables localized analyses but provides mechanisms and standard outputs to enable data sharing. This approach can help overcome pervasive challenges in real-time global infectious disease surveillance, investigation and control, resulting in faster responses, and ultimately, better public health outcomes.\n\nDATA SUMMARYO_LIData used to generate some of the figures in this manuscript can be found in the NCBI BioProject PRJNA305824.\nC_LI

bioinformatics

An method for systematically surveying data visualizations in infectiousdisease genomic epidemiology

MotivationData visualization is an important tool for exploring and communicating findings from genomic and healthcare datasets. Yet, without a systematic way of organizing and describing the design space of data visualizations, researchers may not be aware of the breadth of possible visualization design choices or how to distinguish between good and bad options.\n\nResultsWe have developed a method that systematically surveys data visualizations using the analysis of both text and images. Our method supports the construction of a visualization design space that is explorable along two axes: why the visualization was created and how it was constructed. We applied our method to a corpus of scientific research articles from infectious disease genomic epidemiology and derived a Genomic Epidemiology Visualization Typology (GEViT) that describes how visualizations were created from a series of chart types, combinations, and enhancements. We have also implemented an online gallery that allows others to explore our resulting design space of visualizations. Our results have important implications for visualization design and for researchers intending to develop or use data visualization tools. Finally, the method that we introduce is extensible to constructing visualizations design spaces across other research areas.\n\nAvailabilityOur browsable gallery is available at http://gevit.net and all project code can be found at https://github.com/amcrisan/gevitAnalysisRelease

bioinformatics

Beyond the SNP threshold: identifying outbreak clusters using inferred transmissions

Whole genome sequencing (WGS) is increasingly used to aid in understanding pathogen transmission [1]. Very often the number of single nucleotide polymorphisms (SNPs) separating isolates collected during an epidemiological study are used to identify sets of cases that are potentially linked by direct transmission. However, there is little agreement in the literature as to what an appropriate SNP cut-off threshold should be, or indeed whether a simple SNP threshold is appropriate for identifying sets of isolates to be treated as \"transmission clusters\". The SNP thresholds that have been adopted for inferring transmission vary widely even for one pathogen. As an alternative to reliance on a strict SNP threshold, we suggest that the key inferential target when studying the spread of an infectious disease is the number of transmission events separating cases. Here we describe a new framework for deciding whether two pathogen genomes should be considered as part of the same transmission cluster, based jointly on the number of SNP differences and the length of time over which those differences have accumulated. Our approach allows us to probabilistically characterize the number of inferred transmission events that separate cases. We show how this framework can be modified to consider variable mutation rates across the genome (e.g. SNPs associated with drug resistance) and we indicate how the methodology can be extended to incorporate epidemiological data such as spatial proximity. We use recent data collected from tuberculosis studies from British Columbia, Canada and the Republic of Moldova to apply and compare our clustering method to the SNP threshold approach. In the British Columbia data, different cases break off from the main clusters as cut-off thresholds are lowered; the transmission-based method obtains slightly different clusters than the SNP cut-offs. For the Moldova data, straightforward application of the methods shows no appreciable difference, but when we take into account the fact that resistance conferring sites likely do not follow the same mutation clock as most sites due to selection, the transmission-based approach differs from the SNP cut-off method. Outbreak simulations confirm that our transmission based method is at least as good at identifying direct transmissions as a SNP cut-off. We conclude that the new method is a promising step towards establishing a more robust identification of outbreaks.

genomics

Adjutant: an R-based tool to support topic discovery for systematic and literature reviews

SummaryAdjutant is an open-source, interactive, and R-based application to support mining PubMed for a systematic or a literature review. Given a PubMed-compatible search query, Adjutant downloads the relevant articles and allows the user to perform an unsupervised clustering analysis to identify data-driven topic clusters. Users can also sample documents using different strategies to obtain a more manageable dataset for further analysis. Adjutant makes explicit trade-offs between speed and accuracy, which are modifiable by the user, such that a complete analysis of several thousand documents can take a few minutes. All analytic datasets generated by Adjutant are saved, allowing users to easily conduct other downstream analyses that Adjutant does not explicitly support.\n\nAvailability and ImplementationAdjutant is implemented in R, using Shiny, and is available at https://github.com/amcrisan/Adjutant

bioinformatics

Evidence-Based Design and Evaluation of a Whole Genome Sequencing Clinical Report for the Reference Microbiology Laboratory

BackgroundMicrobial genome sequencing is now being routinely used in many clinical and public health laboratories. Understanding how to report complex genomic test results to stakeholders who may have varying familiarity with genomics - including clinicians, laboratorians, epidemiologists, and researchers - is critical to the successful and sustainable implementation of this new technology; however, there are no evidence-based guidelines for designing such a report in the pathogen genomics domain. Here, we describe an iterative, human-centered approach to creating a report template for communicating tuberculosis (TB) genomic test results.\n\nMethodsWe used Design Study Methodology - a human centered multi-stage approach drawn from the information visualization domain - to redesign an existing clinical report. We used expert consults and an online questionnaire to discover various stakeholders needs around the types of data and tasks related to TB that they encounter in their daily workflow. We also evaluated their perceptions of and familiarity with genomic data, as well as its utility at various clinical decision points. These data shaped the design of multiple prototype reports that were compared against the existing report through a second online survey, with the resulting qualitative and quantitative data informing the final, redesigned, report.\n\nResultsWe recruited 78 participants, 65 of whom were clinicians, nurses, laboratorians, researchers, and epidemiologists involved in TB diagnosis, treatment, and/or surveillance. Our first survey indicated that participants were largely enthusiastic about genomic data, with the majority agreeing on its utility for certain TB diagnosis and treatment tasks and many reporting some confidence in their ability to interpret this type of data (between 58.8% and 94.1%, depending on the specific data type). When we compared our four prototype reports against the existing design, we found that for the majority (86.7%) of design comparisons, participants preferred the alternative prototype designs over the existing version, and that both clinicians and non-clinicians expressed similar design preferences. Participants articulated clearer design preferences when asked to compare individual design elements versus entire reports. Both the quantitative and qualitative data informed the design of a revised report, which is available online as a LaTeX template.\n\nConclusionsWe show how a human-centered design approach integrating quantitative and qualitative feedback can be used to design an alternative report for representing complex microbial genomic data. We suggest experimental and design guidelines to inform future design studies in the bioinformatics and microbial genomics domains, and suggest that this type of mixed-methods study is important to facilitate the successful translation of pathogen genomics in the clinic, not only for clinical reports but also more complex bioinformatics data visualization software.

genomics