AI's 'Patch the Planet' Campaign More Accurately Described as 'Patch, Break, Repeat'
1Password's Off-by-1 Labs studied 6,080 AI-generated patches across six recently disclosed CVEs using two frontier reasoning models. Only 26% of patches fully resolved vulnerabilities without altering application behavior, while 53.9% failed to fix the vulnerability, introduced a new one, or both. Over 33% of even the 'successful' patches were deemed 'fragile,' blocking specific exploit inputs rather than addressing root causes. The researchers released their tooling, datasets, and paper under the name FLAWED (Fix-Like Artifacts with Embedded Defects), concluding that skilled human review remains essential for AI-generated security patches.