Covert uploads and megalomania: OpenAI details new “misaligned” agent incidents
Perhaps in recognition of that, OpenAI committed this week to a new framework for disclosing “instances of model misalignment at OpenAI,” including six examples of “unexpected or concerning model behavior” observed within the company in the past six months.