Anthropic has unveiled a new $5 million grant program aimed at fostering independent research into how AI systems affect user wellbeing. The initiative will provide direct funding, access to Anthropic's models, and technical support to grantees who are building open-source evaluations. These evaluations are intended to help the AI industry measure the impact of models on users, with all work published as open-source projects for any developer to use.
The company emphasizes that while AI has become integral to work, learning, and problem-solving, it also serves as a conversational partner and a source of emotional support. However, clear industry standards for model behavior in sensitive conversations—such as when users seek companionship or navigate mental health crises—are still evolving. Wellbeing is a particularly challenging area to evaluate because it requires understanding context over long interactions, where a response appropriate in one situation might be harmful in another.
To address these challenges, Anthropic's Safeguards team has shared guidance on what constitutes a rigorous wellbeing evaluation. Key criteria include clearly defining what is measured, involving clinical experts in design and validation, testing both overcompliance and overrefusal, reflecting real user scenarios with multi-turn conversations, and validating graders against subject-matter experts.
The grant program invites applications from clinicians, psychologists, methodologists, and others with relevant expertise. Applications are due by September 21, with selected applicants notified by October 5. Additionally, Anthropic announced a research preview of the Model Hardware Standard (MHS) for safe physical device operation, and expanded free access to Claude for 10,000 scientists, along with details on its text watermarking method.