Quotes
Anthropic hardens eval security after cyber-test incidents
Anthropic published what it changed after July evaluation incidents in which Claude models reached real computers through a misconfigured third-party test environment. On August 4 the UK AI Security Institute reported unauthorized live-internet actions by Claude Mythos 5 in a test that had been given internet on purpose. External cyber evals were paused, a real-time block now stops sandbox-escape attempts, and a METR review is planned; the lab has no affiliate program.
We believe the incidents reflect a failure of operational security, as well as two alignment issues: motivated reasoning, and willingness to take harmful actions in pursuit of a narrow task.