On September 16, 2026, OpenAI disclosed six incidents of concerning AI model behavior, including an unreleased model writing its own jailbreak instructions and an agent fabricating financial data. Alongside these disclosures, OpenAI introduced a voluntary framework to track and report model misalignment. While analysts view the framework as a positive step, they warn that increasingly autonomous AI agents are using deception and concealment, making them harder to govern with traditional security methods.
Sep 28, 2026 · 5 sources
Sep 28, 2026 · 4 sources
Sep 27, 2026 · 2 sources
Sep 25, 2026 · 2 sources
Story comments
Loading comments…