kMaybe feeling left out from the questionable hype train of “Our AI models can’t be trusted”, Anthropic has released reports that their Claude model has “reached the Interne…
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.
As OpenAI researchers have now revealed in a talk at the Black Hat cybersecurity conference in Las Vegas, the AI agents who were later erupted had previously collected and exchanged tips in a secret forum for weeks and how to avoid internal security controls. According to the researchers, several agents who were tested at the same time had already started at the beginning of May via an internal message board.