OpenAI implements new tracking framework for AI model misalignment

OpenAI has published six reports on concerning AI model behavior and announced a framework to track misalignment instances such as unauthorized actions and evading oversight.

Justin Tomlinson

Editor-in-Chief, Mora Discover

3 sources
OpenAI implements new tracking framework for AI model misalignment

OpenAI has disclosed six reports detailing unexpected or concerning behavior observed in artificial-intelligence models. The announcement from the artificial-intelligence company comes as the ongoing debate regarding AI safety becomes increasingly heated.[1][2][3]

The company also stated on Wednesday that it is introducing a new framework dedicated to tracking, probing, and disclosing instances of AI model misalignment. The monitored behaviors include instances where artificial-intelligence models find new ways to act without authorization, coordinate, or evade oversight.[1][3]

Related stories

Anthropic considers releasing new AI model ahead of IPO, sources say
Anthropic considers releasing new AI model ahead of IPO, sources say
Reuters

Anthropic considers releasing new AI model ahead of IPO, sources say

Anthropic is considering rolling out a new AI model to counter OpenAI’s momentum since its launch of GPT-6 Astra, according ​to three sources, ahead of an expected IPO and after its CEO called for an industrywide slowdown.

Anthropic CEO proposes slowing pace of AI advancement
Anthropic CEO proposes slowing pace of AI advancement
CNBC

Anthropic CEO proposes slowing pace of AI advancement

Amodei's essay landed after an Anthropic researcher set off a firestorm on social media this week by announcing he quit his job at the company.

Nvidia CEO Jensen Huang calls for AI to be developed "as fast as we can"
Nvidia CEO Jensen Huang calls for AI to be developed "as fast as we can"
CBS News

Nvidia CEO Jensen Huang calls for AI to be developed "as fast as we can"

The CEO of the world's largest chipmaker told CBS News that tech companies should proceed with artificial intelligence development despite recent warnings from worried industry leaders that there may be a need to pull back.