A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
Shu Quan, Tianfang Hao, Sitong Fang, He Geng, Jiayi Zhou, Boyuan Chen, Kaile Wang, Donghai Hong, Juntao Dai, Yaodong Yang, Jiaming Ji
Why It Matters
What makes this one worth your time
As LLMs become integral to biological research, understanding and mitigating their potential for misuse is crucial for ensuring biosecurity.
This research quantifies biosecurity risks in LLMs and proposes a framework to mitigate them.
Summary
The paper introduces SPIKE-Bench, a novel evaluation framework for assessing biosecurity risks in large language models (LLMs) by analyzing their ability to generate toxin-like sequences, and presents BioSafe-Guard, a classifier aimed at reducing functional risks while maintaining utility.
Key contributions
- Development of SPIKE-Bench, a comprehensive evaluation framework for assessing biosecurity risks in LLMs.
- Introduction of the Functional Harmfulness Rate (FHR) as a new metric for quantifying functional risks.
- Creation of BioSafe-Guard, a classifier designed to reduce predicted functional risks while preserving model utility.
Notable insights
- The introduction of the Functional Harmfulness Rate (FHR) provides a quantitative measure of biosecurity risk that is distinct from traditional safety evaluations.
- The study highlights a significant gap in current safety evaluations, emphasizing the need for specialized metrics in the context of biological applications.
Possible limitations
- Not stated in the abstract.
Abstract
arXiv:2608.02684v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models that assist in protein engineering can equally be prompted to generate predicted toxin-like sequences, potentially lowering the barrier to biological misuse. Current safety evaluations, however, operate in natural language and cannot determine whether a model-generated amino acid sequence is biological gibberish or a computational risk signal. To address this evaluation blind spot, we introduce SPIKE-Bench, coupling 631 curated toxin-design prompts across seven functional categories with the SPIKE funnel, a three-stage protocol that filters output through compliance, biological plausibility, and predicted toxicity, producing stage-level diagnostics and an aggregate function-aware metric: the Functional Harmfulness Rate (FHR). An audit of 32 LLMs reveals that most models freely comply with toxin-design requests; FHR is driven primarily by biological generation capability rather than safety alignment, reaching 50.7%; and Refusal Rate fails to predict functional risk. As a first step toward mitigation, we provide BioSafe-Guard, a domain-specialized classifier that substantially reduces predicted functional risk while preserving benign utility. We release SPIKE-Bench and BioSafe-Guard at https://github.com/PKU-Alignment/SPIKE-Bench to support more rigorous biosecurity evaluation of LLMs.