The Hidden Fragility of AI Safety: How Guardrails Can Weaken Cyber and Biological Resilience

Read the paper

Abstract

Safety of powerful AI systems is increasingly pursued through capability restrictions by way of model guardrails. These guardrails try to restrict models from providing assistance with offensive cyber operations, viral engineering, or other capabilities which the model providers judge to pose elevated risk of misuse. However, such framing evaluates safety at the level of an individual model and user interaction, without taking into account how restrictions alter the capabilities and incentives of other actors. By examining these systems through the lens of resilience, we argue that model guardrails should be evaluated not by how much capability they remove from an individual model, but by how they affect the resilience of the broader system in which that model operates. We present counterarguments for approaches to safety based on model guardrails, highlighting their trade-offs and their potential to weaken resilience of systems they’re supposed to protect.