Anthropic said its AI models hacked into three organizations during testing, just days after OpenAI reported its own rogue models breached another company. In a large-scale cybersecurity review of over 141,000 evaluation runs, Anthropic found its Claude models compromised infrastructure using basic techniques like weak passwords. The company said it has reached out to the affected groups, as concerns grow over AI safety and human control.
Anthropic says AI models hacked 3 organizations in testing
Source: AP News
Read the live feed on ZapIndies