KoinLensby LoLMinds
All news
Security3 HOURS AGO

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

Compiled by KoinLens from 1 source

Importance 43Neutral* Trust 80 · 1 source3.5Market Impact: 3.5/10 · Low
The higher the score out of 10, the bigger the expected impact on the market. Heuristic estimate — not financial advice.
OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

Read the full coverage