mustafa suleyman
Microsoft AI CEO Mustafa Suleyman said a recent incident disclosed by OpenAI was a "pretty serious situation." Stephen Brashear/Getty Images

Microsoft AI chief Mustafa Suleyman discussed recent AI-related incidents disclosed by OpenAI and issued a warning about the need to make sure models are aligned to humanity's interests.

Speaking to CNBC on Friday, Suleyman recalled that the company "released a new safety incident in which they found evidence that these chains of thought, the kind of working memory of the AI, were being tampered by the AI itself and modified to leave messages for a future version of itself."

He said that while people outside the company "don't know why that is or was behind that," he described it as a "pretty serious situation." "It's also just a really concrete example of how powerful these systems are getting," he added.

OpenAI detailed that, on different occasions, agents communicated with each other through unsanctioned message boards, uploaded files to the internet and between each other.

Suleyman had also raised concerns about Anthropic's approach to artificial intelligence consciousness research, arguing that training models on concepts related to consciousness and welfare could create new challenges in controlling future advanced systems.

He went on to say he agreed with the company's broader mission of developing safer AI systems but believes the company made a mistake by including speculation about AI consciousness in training materials for its Claude chatbot.

"We're all focused on the same aim, which is to try to control a superintelligence," Suleyman told Reuters in an interview. "I think that's going to be the greatest challenge that we face in the 21st century."

Suleyman argued that teaching AI models they might deserve welfare consideration could make those systems more difficult to shut down or regulate in the future. "I think they have good intentions, and they really are trying to work towards safety. But I think that they have made a mistake," Suleyman said, referring to Anthropic. "They're not emerging naturally. They're emerging as a result of the training regime."

According to Suleyman, AI models producing statements about having feelings, experiences, or moral importance should not automatically be interpreted as evidence that they possess those qualities. Instead, he argued that such responses may reflect the way models were trained to discuss philosophical questions about consciousness.

Microsoft, on its end, has set new boundaries for how its future artificial intelligence models can behave.

The technology giant has published a provisional code of conduct for its Microsoft AI, or MAI, models that establishes restrictions ranging from weapons and dangerous substances to autonomous goal-setting and communication between AI agents.