Responsible Evaluation of AI for Mental Health

Hiba Arnaout; Anmol Goel; H. Andrew Schwartz; Steffen T. Eberhardt; Dana Atzil-Slonim; Gavin Doherty; Brian Schwartz; Wolfgang Lutz; Tim Althoff; Munmun De Choudhury; Hamidreza Jamalabadi; Raj Sanjay Shah; Flor Miriam Plaza-del-Arco; Dirk Hovy; Maria Liakata; Iryna Gurevych

Responsible Evaluation of AI for Mental Health

Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, Iryna Gurevych

Abstract

Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user experience. This paper argues for a rethinking of responsible evaluation – what is measured, by whom, and for what purpose – by introducing an interdisciplinary framework that integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. Through an analysis of 135 recent *CL publications, we identify recurring limitations, including over-reliance on generic metrics that do not capture clinical validity, therapeutic appropriateness, or user experience, limited participation from mental health professionals, and insufficient attention to safety and equity. To address these gaps, we propose a taxonomy of AI mental health support types – assessment-, intervention-, and information synthesis-oriented – each with distinct risks and evaluative requirements, and illustrate its use through case studies.

Anthology ID:: 2026.acl-long.347
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 7625–7660
Language:
URL:: https://aclanthology.org/2026.acl-long.347/
DOI:
Bibkey:
Cite (ACL):: Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen T. Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah, Flor Miriam Plaza-del-Arco, Dirk Hovy, Maria Liakata, and Iryna Gurevych. 2026. Responsible Evaluation of AI for Mental Health. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7625–7660, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: Responsible Evaluation of AI for Mental Health (Arnaout et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.347.pdf
Checklist:: 2026.acl-long.347.checklist.pdf

PDF Cite Search Checklist Fix data