When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

Mohammad Alijanpour Shalmani, Alale Rezvani Boroujeni, Jiann S. Yuan


Abstract
Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, or surface mismatched information, inventing venues, confirmations, or booking details not grounded in the database. We study a lightweight prompting-based recovery approach that improves robustness without retraining or additional model calls. We compare three response strategies, including a guided recovery prompt conditioned on structured database status, across six open-weight model families (DeepSeek-R1, Gemma-2, Llama-3, Mistral, Phi-3, and Qwen-2.5) and four database conditions: empty result, wrong-domain retrieval, API error, and clean retrieval. Using fault-injected benchmarks built on two structurally different datasets, MultiWOZ 2.2 (5 domains) and SGD (20 domains), we find that naive agents hallucinate on 30.5% of failure turns on MultiWOZ and 20.9% on SGD. Our Guided-Retry strategy reduces hallucination by 50% on MultiWOZ (30.5→15.3%) and by 42% on SGD (20.9→12.2%) without retraining. However, residual hallucination remains substantial (6–37% across models), with wrong-domain failures the hardest case. Results are consistent across both datasets and all six model families, and human annotation shows substantial agreement while supporting the validity of the automatic commitment-safety metric.
Anthology ID:
2026.sigdial-1.57
Volume:
Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue
Month:
August
Year:
2026
Address:
Atlanta, Georgia, USA
Editors:
Jinho D. Choi, Yun-Nung Chen, Kotaro Funakoshi, Ali Emami
Venue:
SIGDIAL
SIG:
SIGDIAL
Publisher:
Association for Computational Linguistics
Note:
Pages:
815–820
Language:
URL:
https://aclanthology.org/2026.sigdial-1.57/
DOI:
Bibkey:
Cite (ACL):
Mohammad Alijanpour Shalmani, Alale Rezvani Boroujeni, and Jiann S. Yuan. 2026. When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue. In Proceedings of the 27th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 815–820, Atlanta, Georgia, USA. Association for Computational Linguistics.
Cite (Informal):
When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue (Alijanpour Shalmani et al., SIGDIAL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.sigdial-1.57.pdf