Vetting practitioners beats blunt safety filters
Frontier model guardrails create persistent friction for security teams. The prompt patterns required to dissect malware or trace an exploit chain look identical to attack payloads, which means default API filters routinely refuse the work defenders need done. Model providers have spent years trying to tune refusals around intent, but natural language intent is notoriously difficult to classify at the inference boundary.
Anthropic is shifting that boundary from the prompt to the user. By expanding its Cyber Verification Program, verified security professionals can access Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 under specialized safeguards, including dedicated tiers for authorized penetration testing and red-teaming. Vetting the practitioner instead of over-censoring the model makes technical sense. You cannot run a competent red team or patch subtle software vulnerabilities with a model that refuses to explain how an exploit executes.
The trade-off is organizational. Once safety enforcement relies on identity checks, the critical attack surface becomes the vetting pipeline and API key security of authorized accounts rather than prompt evasion. For enterprise security teams, access to uncrippled reasoning on frontier models will be valuable, provided labs can keep the credentialing process fast enough to match active development cycles.