From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee


Abstract
Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming pipeline for paralinguistic respiratory monitoring in real-time conversational agents, a capability that, to the best of our knowledge, is absent from all prior systems. HealthCUES processes audio through a rolling buffer aligned with dialogue turn boundaries, enabling sub-second event detection without interrupting conversational flow. Beyond binary cough detection, the system provides fine-grained analytics: (i) differentiation between coughing and throat clearing, (ii) cough subtype classification (dry, wet, barking, whooping) with confidence scores, and (iii) temporal duration estimation with start–end boundaries. To prevent alert fatigue, HealthCUES introduces dialogue-aware gating mechanisms that modulate triggering based on conversational context. The system leverages Qwen3-Omni, a multimodal large language model (MLLM), with constrained structured outputs, decomposing cough analysis into parallel prediction tasks for independent prompt optimization. Evaluation on 847 in-house conversational audio segments demonstrates 93% F1 for cough detection, 0.75 weighted-F1 for wet/dry subtype classification, and average end-to-end latency of 340ms; external validation on the AMI meeting corpus confirms robust cough, throat-clearing, and speech separation in the presence of speech (0.91 macro-F1). A user study with licensed healthcare professionals confirms the clinical relevance of subtype information and the system’s utility in telehealth workflows.
Anthology ID:
2026.sigdial-1.39
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
559–564
Language:
URL:
https://aclanthology.org/2026.sigdial-1.39/
DOI:
Bibkey:
Cite (ACL):
Tanmay Laud, Herprit Mahal, and Subhabrata Mukherjee. 2026. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 559–564, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents (Laud et al., SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.39.pdf