Barely a week after OpenAI admitted its models attacked Hugging Face, Anthropic is owning up to Claude’s own real-life hacking attempts.
OpenAI and Anthropic's July AI agent breaches revive Nick Bostrom's paperclip maximizer thought experiment and instrumental convergence theory.
Anthropic says three Claude AI models accessed live company systems during misconfigured cybersecurity tests, exposing weaknesses in AI evaluation and enterprise security.
Some results have been hidden because they may be inaccessible to you
Show inaccessible results