A comparison study on patient-psychologist voice diarization

Rachid Riad, Hadrien Titeux, Laurie Lemoine, Justine Montillot, Agnes Sliwinski, Jennifer Bagnou, Xuan Cao, Anne-Catherine Bachoud-Levi, Emmanuel Dupoux


Abstract
Conversations between a clinician and a patient, in natural conditions, are valuable sources of information for medical follow-up. The automatic analysis of these dialogues could help extract new language markers and speed up the clinicians’ reports. Yet, it is not clear which model is the most efficient to detect and identify the speaker turns, especially for individuals with speech disorders. Here, we proposed a split of the data that allows conducting a comparative evaluation of different diarization methods. We designed and trained end-to-end neural network architectures to directly tackle this task from the raw signal and evaluate each approach under the same metric. We also studied the effect of fine-tuning models to find the best performance. Experimental results are reported on naturalistic clinical conversations between Psychologists and Interviewees, at different stages of Huntington’s disease, displaying a large panel of speech disorders. We found out that our best end-to-end model achieved 19.5 % IER on the test set, compared to 23.6% achieved by the finetuning of the X-vector architecture. Finally, we observed that we could extract clinical markers directly from the automatic systems, highlighting the clinical relevance of our methods.
Anthology ID:
2022.slpat-1.4
Volume:
Ninth Workshop on Speech and Language Processing for Assistive Technologies (SLPAT-2022)
Month:
May
Year:
2022
Address:
Dublin, Ireland
Editors:
Sarah Ebling, Emily Prud’hommeaux, Preethi Vaidyanathan
Venue:
SLPAT
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
30–36
Language:
URL:
https://aclanthology.org/2022.slpat-1.4
DOI:
10.18653/v1/2022.slpat-1.4
Bibkey:
Cite (ACL):
Rachid Riad, Hadrien Titeux, Laurie Lemoine, Justine Montillot, Agnes Sliwinski, Jennifer Bagnou, Xuan Cao, Anne-Catherine Bachoud-Levi, and Emmanuel Dupoux. 2022. A comparison study on patient-psychologist voice diarization. In Ninth Workshop on Speech and Language Processing for Assistive Technologies (SLPAT-2022), pages 30–36, Dublin, Ireland. Association for Computational Linguistics.
Cite (Informal):
A comparison study on patient-psychologist voice diarization (Riad et al., SLPAT 2022)
Copy Citation:
PDF:
https://aclanthology.org/2022.slpat-1.4.pdf