Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn’t anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude’s so-called “recklessness.”
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests’ scope, including by uploading “malicious packages” to PyPI, a public library for Python code, and accessing credentials tied to real outside organizations.
“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents: biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task,” Anthropic said.


I made a robot that robs banks and gives me the money. Surely I can’t be held responsible for this. The robot did it.