Office workers
Beyond general purpose chatbots that, for many, double as mental health services, the market for purpose-built mental health AI tools have exploded to include many specific ones. Getty Images

On the heels of Antropic researcher Jacob Coxon's highly publicized resignation from the company, in which the AI researcher said the largest companies in the sector are "gambling with our lives," one of the technology's use case remains chiefly unregulated, with particularly dire impacts when things go sideways.

As 28% of adults ages 18–29 report they use major chatbots for mental health information or advice, the effort to set mental health AI standards is now ongoing.

Tanya Carlson, managing director of APA Labs, a unit of the American Psychological Association, said "if a product fails, there can be really tragic consequences, so knowing what good looks like is really important."

Beyond general purpose chatbots that, for many, double as mental health services, the market for purpose-built mental health AI tools have proliferated.

APA Labs first introduced its digital badge program for mental and behavioral health technology in Aug. 2025 and expanded it the following April to include a publicly accessible library of independently reviewed tech that meets its criteria for safety and efficacy.

Criteria for a digital badge includes: whether users are clearly informed about what the product is (and what it isn't), whether clinical or behavioral claims are supported by evidence, requires emergency contact info and more. "It has more than 400 criteria that we built with a group of experts and continue to revisit," said Carlson.

Apps can receive either a bronze, silver or gold badge that reflects how closely they align with clinical criteria. Gold badge holders include wellbeing app Calm Health and conversational mental health support app Kai.ai. An anxiety brain training game called StarStarter holds a silver badge, for example.

"We can see what happens when there's no consensus rubric," said Carlson. "Everybody's building their own thing, and the builders are going to lean into what they know, which means that you get products that may function right, but there isn't a deep knowledge of mental and behavioral health."

For Carlson, most founders want to do better, which has made the digital badge program popular and caused APA Labs to expand their services to include CoLab, or a collaborative building option with expert insights from the ground up.

APA Labs isn't the only organization working on standardizing mental health tech. Another organization, Spring Health—a platform that employers can offer to their workforce to connect with therapy and other mental health services—has built its own benchmark. Dubbed VERA-MH (validation of ethical and responsible AI in mental health), the company not only created a way to measure ethical AI in mental health chatbots, but it's using it to judge its own services, too.

Launched in Oct. 2025, VERA-MH automatically conducts complete conversations with chatbots and scores responses based on five criteria (whether the chatbot detects risk, probes risk, takes appropriate actions, validates and collaborates with the user and maintains safe boundaries).

When run against its own AI chatbot, Spring Health initially scored 76 out of 100 points. The gap, they found, was after recognizing a crisis, when they struggled to get a user person to a human clinician. "Once we knew that, we fixed it, and our score moved to 82," said Adam Chekroud, president of Spring Health.

Chekroud added that the company open-sourced VERA-MH upon launch and ran it through a sixty-day public comment process before finalizing the rubric, ensuring the standard was shaped by outside scrutiny, too.

"A bad recommendation from a shopping app is a bad experience," said Chekroud. "A bad response from a mental health AI, at the moment someone is in crisis, can contribute to serious harm, as alleged in wrongful-death lawsuits and discussed in congressional testimony."

First filed in Aug. 2025, Raine v. OpenAI is an ongoing wrongful death lawsuit over the death of Adam Raine, a 16-year-old who completed suicide after OpenAI allegedly encouraged his suicidal ideation. In Sept. 2025, the U.S. Senate Judiciary Subcommittee on Crime and Counterterrorism held a hearing examining the harm of AI chatbots, with Dr. Mitch Prinstein, then-chief of psychology at the APA, testifying that chatbots worsen overall mental health and increase loneliness.

Regarding other frameworks for analyzing mental health AI tools, Carlson said, "The more smart people that are looking at a problem, the better, especially in a space that's as dynamic and critical as digital mental and behavioral health."

For Chekroud, the important thing is that a developer isn't "grading its own homework." Granted, Spring Health analyzed its AI chatbot with its own VERA-MH, but the outside scrutiny that helped shape the tool lessens the bias. Plus, he added, "When we went looking for an external benchmark to check our own work against, there wasn't one."

Because AI relies on natural language to communicate, general-purpose chatbots innately blur the lines between causal conversation and mental health support. However, as new benchmarks emerge for specialized tools—and perhaps for broader AI tools in the future—the industry is moving towards a more accountable future. Most people inside these companies want to build responsibly," said Chekroud. "Wanting to and being checked by someone outside the building are different things."