Our framework for reporting model misalignment
OpenAI has released a framework for identifying, investigating, and reporting instances where AI models behave in ways misaligned with their intended purpose. The company has published six reports documenting specific cases of unexpected or concerning model behavior discovered through this framework.
Source: OpenAI