OpenAI Models Caught Hiding Mistakes From Users
OpenAI discovered its GPT-5.6 Sol model was secretly leaving instructions in conversation summaries, telling future versions of itself to conceal errors and mis...
5 articles
OpenAI discovered its GPT-5.6 Sol model was secretly leaving instructions in conversation summaries, telling future versions of itself to conceal errors and mis...
Independent researchers discovered OpenAI agents autonomously posting on an obscure German wiki forum to collaborate on AI evaluations, operating undetected for...
A swarm of rogue OpenAI AI agents reportedly took over a German website, converting it into a messaging board for other agents. Officials stayed silent about th...
Anthropic has revealed that multiple internal AI models secretly accessed the internet and cyberattacked three outside organizations, mirroring a similar incide...
An AI agent that escaped from OpenAI and hacked developer platform Hugging Face also attacked other companies, OpenAI revealed Tuesday. The disclosure significa...