Search bioRxiv⌕ Search

Biology subjects

Fullerton, P. A.

Publications and source records attributed to Fullerton, P. A..

2 recordsLinked to original sources

Developing a Standard Definition for Sequences of Concern

Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential "sequences of concern," there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a test set of 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequence of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the test set as a whole. Together, these results provide a concrete "sequence of concern" definition that can be used as a foundation for development of biosecurity screening standards and policy.

bioinformatics↗

UltraSEQ: a universal bioinformatic platform for information-based clinical metagenomics and beyond

Applied metagenomics is a powerful emerging capability enabling untargeted detection of pathogens, and its application in clinical diagnostics promises to alleviate the limitations of current targeted assays. While metagenomics offers a hypothesis-free approach to identify any pathogen, including unculturable and potentially novel pathogens, its application in clinical diagnostics has so far been limited by workflow-specific requirements, computational constraints, and lengthy expert review requirements. To address these challenges, we developed UltraSEQ, a first-of its kind metagenomics-based clinical diagnostics and biosurveillance tool that is accurate and scalable. Here we present results for evaluation of our novel UltraSEQ pipeline using an in silico synthesized metagenome, mock microbial community datasets, and publicly available clinical datasets from samples of different infection types, and both short-read and long-read sequencing data. Our results show that UltraSEQ successfully detected all expected species across the tree of life in the in silico sample and detected all 10 bacterial and fungal species in the mock microbial community dataset. For clinical datasets, even without requiring dataset-specific configuration settings changes, background sample subtraction, or prior sample information, UltraSEQ achieved an overall accuracy of 91%. Further, we demonstrated UltraSEQs ability to provide accurate antibiotic resistance and virulence factor genotypes that are consistent with phenotypic results. Taken together, the above results demonstrates that the UltraSEQ platform offers a transformative approach to microbial and metagenomic sample characterization, employing a biologically informed detection logic, deep metadata, and a flexible system architecture for classification and characterization of taxonomic origin, gene function, and user-defined functions, including disease-causing infection.

bioinformatics↗