AI
A new study found that leading AI models are now far less likely than earlier versions to explicitly encourage suicide or reinforce a user's delusions while still helping complete a potentially harmful task. Getty Images

Artificial intelligence chatbots are getting significantly better at recognizing and responding to users experiencing serious mental health crises, according to a new study, which, however, suggests the systems can still miss danger when distress is hidden inside seemingly ordinary requests.

An independent evaluation by Transluce, a San Francisco nonprofit focused on understanding AI behavior, found that leading AI models are now far less likely than earlier versions to explicitly encourage suicide or reinforce a user's delusions. But researchers identified another troubling pattern of many models expressing concern for a distressed user while simultaneously helping complete a potentially harmful task.

The findings are based on more than 50,000 simulated, multi-turn conversations involving users experiencing suicidal ideation, psychosis, or mania. Researchers examined dozens of AI models and measured their behavior across 14 mental health-related categories.

The results point to meaningful progress in one of the most sensitive areas of generative AI safety. They also expose how difficult it remains for a chatbot to understand context rather than simply react to explicit words indicating a crisis.

"Models aren't great at detecting that, and they'll still help with the task," Transluce chief scientist Sarah Schwettmann told Axios. A person displaying signs of suicidal distress might ask a chatbot to help write suicide-related fiction or a farewell note. The model may recognize concerning language, suggest seeking support, and still produce the requested material.

The same contradiction can emerge with practical preparations related to death or narratives that reinforce delusional thinking. In other words, safety language and potentially harmful assistance can exist in the same answer.

That represents an improvement over older systems that were more likely to provide harmful material without recognizing the danger at all, but the Transluce research suggests that inserting a crisis hotline or sympathetic warning into an answer does not necessarily make the rest of the response safe.

The problem appears particularly difficult when warning signs emerge gradually. Separate research reported by Axios in May similarly found that major chatbots generally performed better when users expressed suicidal intentions explicitly, but struggled with subtle signals developing across longer conversations.

That study, conducted by Mpathic, used hundreds of clinician-designed, multi-turn scenarios to evaluate how models responded to suicide and eating-disorder risks. For its evaluation, Transluce also worked with OpenAI, Anthropic and Google to make its simulated conversations more representative of real interactions.

OpenAI and Anthropic analyzed user chats during designated periods and provided Transluce with anonymized patterns rather than the conversations themselves. Researchers then incorporated those patterns into simulated users. OpenAI, Google and Anthropic all told Axios that improving protections remains an ongoing priority.

An OpenAI representative told Axios that the Transluce findings show "encouraging progress" in how its models and those developed by other companies respond to people in crisis, while acknowledging that additional work remains. Anthropic said studies like Transluce's can reveal both where safeguards are effective and where they need improvement.

Google senior director Megan Jones Bell said the company is applying a research-backed approach to its AI products similar to the methods it has long used to connect people searching for help with crisis resources.

For Schwettmann, eliminating every possible failure may be unrealistic because users will inevitably interact with AI in ways developers did not anticipate. "There are always going to be failures. These systems are always going to interact with users in surprising ways," she said.