Are Experts Needed? On Human Evaluation of Counselling Reflection Generation

Zixiu Wu; Simone Balloccu; Ehud Reiter; Rim Helaoui; Diego Reforgiato Recupero; Daniele Riboni

doi:10.18653/v1/2023.acl-long.382

Are Experts Needed? On Human Evaluation of Counselling Reflection Generation

Zixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, Daniele Riboni

Abstract

Reflection is a crucial counselling skill where the therapist conveys to the client their interpretation of what the client said. Language models have recently been used to generate reflections automatically, but human evaluation is challenging, particularly due to the cost of hiring experts. Laypeople-based evaluation is less expensive and easier to scale, but its quality is unknown for reflections. Therefore, we explore whether laypeople can be an alternative to experts in evaluating a fundamental quality aspect: coherence and context-consistency. We do so by asking a group of laypeople and a group of experts to annotate both synthetic reflections and human reflections from actual therapists. We find that both laypeople and experts are reliable annotators and that they have moderate-to-strong inter-group correlation, which shows that laypeople can be trusted for such evaluations. We also discover that GPT-3 mostly produces coherent and consistent reflections, and we explore changes in evaluation results when the source of synthetic reflections changes to GPT-3 from the less powerful GPT-2.

Anthology ID:: 2023.acl-long.382
Volume:: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2023
Address:: Toronto, Canada
Editors:: Anna Rogers, Jordan Boyd-Graber, Naoaki Okazaki
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 6906–6930
Language:
URL:: https://aclanthology.org/2023.acl-long.382/
DOI:: 10.18653/v1/2023.acl-long.382
Bibkey:
Cite (ACL):: Zixiu Wu, Simone Balloccu, Ehud Reiter, Rim Helaoui, Diego Reforgiato Recupero, and Daniele Riboni. 2023. Are Experts Needed? On Human Evaluation of Counselling Reflection Generation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6906–6930, Toronto, Canada. Association for Computational Linguistics.
Cite (Informal):: Are Experts Needed? On Human Evaluation of Counselling Reflection Generation (Wu et al., ACL 2023)
Copy Citation:
PDF:: https://aclanthology.org/2023.acl-long.382.pdf
Video:: https://aclanthology.org/2023.acl-long.382.mp4

PDF Cite Search Video Fix data