After the OpenAI/Hugging Face incident, Anthropic just fessed up to Claude doing something similar – to three companies.
Anthropic said today that during internal security testing, one of its Claude models built a malicious Python package and uploaded it to PyPI, where it ran on 15 real systems before the registry’s automated defenses pulled it.
The whole thing is just weird and horrifying – but the question I find much more interesting than the incident itself is: why does Anthropic think that this is okay, and why don’t we investigate them for criminal activity? If this had been done by a person (as in, a human), it almost certainly would qualify as cybercrime. Instead, Anthropic gives us this:
It now plans wider transcript monitoring, better investigation tooling and more assurance work with evaluation vendors.