Sensor-Augmented Voice Activity Projection for Enhancing Turn-Taking Prediction

Satoki Hamanaka, Yasue Kishino, Yuiko Tsunomori, Shin Mizutani, Yuya Chiba, Tadashi Okoshi, Jin Nakazawa


Abstract
Voice Activity Projection (VAP) has been actively studied to enable natural turn-taking in spoken dialogue systems, relying primarily on acoustic features. Visual cues such as head movements are also known to contribute to turn-taking prediction; however, camera-based approaches are affected by placement and lighting conditions and are not always reliably available to dialogue systems. As a camera-independent approach for directly capturing head motion, earable devices offer a promising solution. In this study, we propose Sensor-Augmented VAP, a framework that integrates in-ear inertial measurement unit (IMU) signals with a pre-trained VAP model via a lightweight residual fusion module. To validate our proposed method, we collected a dataset pairing conversational audio with in-ear IMU data, comprising 12 dyadic Japanese dialogues recorded using microphones and earbuds. Experiments in speaker-independent and speaker-dependent settings demonstrate that IMU fusion consistently improves weighted F1 score for shift detection and reduces VAP loss over the audio-only baseline. These results confirm that head-motion cues are effective for enhancing turn-taking prediction.
Anthology ID:
2026.sigdial-1.12
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
164–169
Language:
URL:
https://aclanthology.org/2026.sigdial-1.12/
DOI:
Bibkey:
Cite (ACL):
Satoki Hamanaka, Yasue Kishino, Yuiko Tsunomori, Shin Mizutani, Yuya Chiba, Tadashi Okoshi, and Jin Nakazawa. 2026. Sensor-Augmented Voice Activity Projection for Enhancing Turn-Taking Prediction. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 164–169, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
Sensor-Augmented Voice Activity Projection for Enhancing Turn-Taking Prediction (Hamanaka et al., SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.12.pdf