Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization

Anna Manaseryan, Casey Kennington


Abstract
Emotion expression is essential for human-robot interaction, yet current systems rely on static models that cannot adapt to individual users. We present an online reinforcement learning framework that adapts robot emotional behavior policy during live dialogue using binary human feedback. The system integrates a DeBERTa-v3-base emotion classifier and applies Group Relative Policy Optimization (GRPO) in a human-robot dialogue system. At each dialogue turn, the classifier samples a group of emotion candidates and the selected emotion is passed to a generative model that synthesizes a novel robot emotional behavior. We evaluate the system in three experiments: (1) offline supervised fine-tuning followed by GRPO on synthetic dialogue data, (2) a live GRPO training with a human teacher and (3) a final experiment with human participants. Results indicate that the robot was perceived as responsive and emotionally consistent, with high ratings for personality coherence and contextual appropriateness of emotional behaviors. Results further show that online GRPO with human feedback enables effective real-time emotion adaptation in embodied interaction.
Anthology ID:
2026.sigdial-1.41
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
585–596
Language:
URL:
https://aclanthology.org/2026.sigdial-1.41/
DOI:
Bibkey:
Cite (ACL):
Anna Manaseryan and Casey Kennington. 2026. Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 585–596, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
Adaptive Emotion Management in Human-robot Dialogue using Online Group Relative Policy Optimization (Manaseryan & Kennington, SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.41.pdf