Anthropic Offers $5M to Study AI Chatbot Safety in Crisis
New grants target geographic and linguistic gaps in mental health AI benchmarks that leave vulnerable users at risk.

Anthropic launches research grants for crisis AI safety
Anthropic is committing $5 million to fund independent research into how AI chatbots affect users in emotional distress, with individual grants ranging from $500,000 to $1.5 million. The initiative targets a critical blind spot: current safety benchmarks are built almost entirely around English-language inputs and Western clinical standards.
According to ICT Works, which first reported the program, the funding will support between three and ten research teams. Recipients will receive API access or credits alongside the cash awards, plus occasional technical consultation from Anthropic staff. All findings must be released as open-source resources, and Anthropic has committed not to veto or delay publication of results.
The geographic safety gap
The stakes are measurable. A June 2025 study by Miles McCain and colleagues at Anthropic analyzed roughly 4.5 million Claude.ai conversations and found 2.9% involved emotional or personal exchanges. That represents tens of thousands of vulnerable interactions—but the study covered only adults and provided no country or language breakdown.
The World Health Organization's Mental Health Atlas 2024 documents the disparity driving people to chatbots: low-income and lower-middle-income countries have a median of 1.1 to 2.4 specialized mental health workers per 100,000 people, compared to 67.2 in high-income nations. Per-capita mental health spending in wealthy countries reaches $65.89 versus under one dollar in the poorest.
Researchers at ELLIS Alicante, led by Adrián Arnaiz-Rodríguez, tested leading models against crisis inputs from 12 datasets and documented generic, location-inappropriate responses. Self-harm scenarios produced the worst results. A helpline number for the wrong country fails the user who needs it most.
What the grants will fund
Anthropic's Safeguards team has identified priority research areas including detection of harmful emotional dependence and measurement of whether chatbot responses prioritize continued engagement over user welfare. The guidance explicitly calls for evaluation across regional and linguistic variation, including slang, coded language, crisis resources, and cultural norms.
Grantees must disclose Anthropic's funding in all published work. Applications are open until September 21, 2026.
Why it matters
AI safety research has concentrated on preventing models from generating harmful content in controlled test environments. This funding acknowledges a harder problem: chatbots already serve as de facto mental health resources for millions of people in countries with minimal professional infrastructure, yet the systems evaluating their safety reflect only a narrow slice of global need. Without research that accounts for linguistic nuance and regional crisis protocols, AI developers are effectively deploying undertested tools to the users who have the fewest alternatives.
Details of the grant program were reported by ICT Works.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
