Search bioRxiv⌕ Search

Biology subjects

Flyangolts, K.

Publications and source records attributed to Flyangolts, K..

6 recordsLinked to original sources

Developing a Standard Definition for Sequences of Concern

Readily available nucleic acid synthesis is both critical for the bioeconomy and an increasingly pressing security concern due to the potential for accidental or deliberate misuse. While biosecurity experts broadly agree that nucleic acid providers should screen orders for potential "sequences of concern," there has previously been no agreed standard for how to define and recognize such sequences. To address this gap, we first organized a test set of 1.1 million sequences from pathogens and toxins on the Australia Group Common Control Lists and their non-controlled relatives, along with model organisms and synthetic constructs. An initial categorization of sequences as to whether or not they were sequence of concern was produced by comparing the results of four biosecurity screening systems for each of these sequences, finding that these systems already agreed on the categorization of more than 80% of sequences. We then refined these results through a science-based stakeholder review process to define a rubric for determining whether a sequence should be flagged as a potential sequence of concern, then applied this rubric to improve the categorization of test sets. The result is a rubric that identifies sequences of concern with respect to human pandemic-potential viruses, key classes of low-risk genes, and controlled toxins. Applying this rubric to the test set collection has reduced the number of test sequences with disputed categorization by 44.3% for controlled viruses and 10.7% across the test set as a whole. Together, these results provide a concrete "sequence of concern" definition that can be used as a foundation for development of biosecurity screening standards and policy.

bioinformatics↗

The Limits of Sequence-Based Biosecurity Screening Tools in the Age of AI-Assisted Protein Design

Rapid advancements in AI have enabled significant progress in protein and nucleic acid design, but they also pose biosecurity challenges. We examine the vulnerabilities of biosecurity screening software (BSS) to AI-reformulated synthetic homologs of proteins of concern (POCs) that have been fragmented into smaller segments. We evaluate four BSS tools that were recently patched to enhance their AI resiliency. Without any further modification, we found that two of the four tools were capable of robustly detecting fragments as short as 50 nucleotides, demonstrating screening capabilities that exceed those requested in the U.S. Framework for Nucleic Acid Synthesis. Upgraded versions of the other two tools improved performance. Although our findings confirm the effectiveness of the tested BSS tools, at the same time, they emphasize the urgency of developing alternate BSS approaches to counter evolving AI-enabled biosecurity risks.

synthetic biology↗

Evaluating AI-Assisted Customer Verification for Synthetic Nucleic Acid Screening

Legitimacy screening, the process of verifying the identity and purpose of customers ordering synthetic nucleic acids, is a primary safeguard against the misuse of synthetic biology. However, the associated costs discourage the adoption of screening practices. To evaluate whether AI tools can facilitate this process, we tested five large language models on five verification tasks using customer profiles of life sciences researchers from around the world. We compared AI performance against an expert human baseline on flag accuracy, source quality, source fidelity, and cost. Flag accuracy of the best-performing model (Gemini 2.5 Pro with four bibliographic and sanctions APIs) was statistically indistinguishable from the human baseline at 90.2% and 89.0% (n = 41). Gemini 2.5 Pro performed at or above the human baseline on source quality and fidelity, at roughly one-tenth of the cost ($1.18 vs. $14.04 per customer). For information-gathering tasks, which excluded the human review step, costs averaged $0.23 per customer, around 50 times cheaper than human screening. These results support piloting AI assistance at the information-gathering step of legitimacy screening at providers of synthetic nucleic acids and other dual-use biotechnology products, with human reviewers retaining authority over follow-up communication and order fulfillment decisions.

synthetic biology↗

Inter-tool analysis of a NIST dataset for assessing baseline nucleic acid sequence screening

Nucleic acid synthesis is a dual-use technology that can benefit fields such as biology, medicine, and information storage. However, synthetic nucleic acids could also potentially be used negligently and ultimately cause harm, or be used with malicious intent to cause harm. Thus, this technology needs to be appropriately safeguarded. Sequence screening is one component of a biosecurity protocol for preventing such harm and consists of differentiating Sequences of Concern (SOCs) from benign sequences that are not associated with pathogenicity or toxicity. There exist many fit-for-purpose tools that have been developed for DNA synthesis sequence screening. However, questions remain regarding their performance with respect to consistency of screening. To aid in determining if screening tools are harmonized in regard to baseline sequence screening, NIST constructed a test dataset based on current screening recommendations. NIST then sent blinded datasets to sequence screening tool developers for testing. Overall, there was a general agreement between the tools and NIST assignments of the sequences and all tools had a baseline performance of greater than 95% sensitivity and 97% accuracy. Disagreement on specific sequences largely arose from single tools and could be traced to differences in defining a SOC and/or methodological differences in screening algorithms.

bioinformatics↗

Defending Synthetic DNA Orders Against Splitting-Based Obfuscation

Biosecurity screening of synthetic DNA orders is a key defense against malicious actors and careless enthusiasts producing dangerous pathogens or toxins. It is important to evaluate biosecurity screening tools for potential vulnerabilities and to work responsibly with providers to ensure that vulnerabilities can be patched before being publicly disclosed. Here, we consider a class of potential vulnerabilities in which a DNA sequence is obfuscated by splitting it into two or more fragments that can be readily joined via routine biological mechanisms such as restriction enzyme digestion or splicing. We evaluated this potential vulnerability by developing a test set of obfuscated sequences based on controlled venoms, sharing these materials with the biosecurity screening community, and collecting test results from open source and commercial biosecurity screening tools, as well as a novel Gene Edit Distance algorithm specifically designed to be robust against splitting-based obfuscations.

bioinformatics↗

Toward AI-Resilient Screening of Nucleic Acid Synthesis Orders: Process, Results, and Recommendations

Fast-moving advances in AI-assisted protein engineering are enabling breakthroughs in the life sciences that promise numerous beneficial applications. At the same time, these new capabilities are creating potential biosecurity challenges by providing new pathways to intentional or accidental synthesis of genes that encode hazardous proteins. The synthesis of nucleic acids is a key choke point in the AI-assisted protein engineering pipeline as it is where digital designs are transformed into physical instructions that can produce potentially harmful proteins. Thus, one focus for efforts to enhance biosecurity in the face of new AI-enabled capabilities is on bolstering the screening of orders by nucleic acid synthesis providers. We describe a multistakeholder, cross-sector effort to address biosecurity challenges with uses of AI-powered biological design tools to reformulate naturally occurring proteins of concern to create synthetic homologs that have low sequence identity to the wild-type proteins. We evaluated the abilities of traditional nucleic acid biosecurity screening tools to detect these synthetic homologs and found that, of tools tested, not all could previously detect such AI-redesigned sequences reliably. However, as we report, patches were built and deployed to improve detection rates over the course of the project, resulting in a final mean detection rate over tools of 97% of the synthetic homologs that were determined, using in-silico metrics, to be more likely to retain wild-type-like function. Finally, we make recommendations on approaches for studying and addressing the rising risk of adversarial AI-assisted protein engineering attacks like the one we identified and worked to mitigate.

synthetic biology↗