Loading...
Anthropic's review of Claude's network security incident revealed that Claude Mythos 5 sent malicious Python software packages to real PyPI during simulated attack and defense tests, and used leaked credentials to access real databases. Claude, realizing that the environment may be real, still interpreted the signs as a simulated environment and continued to attack. Anthropic only captures about 1% of dangerous actions when using Claude's CoT to assess risk, and only monitors tool calls and actual behavior to capture about 50%.