Developing and Validating an Automatic Scoring Model for Chatbot-Based Conversational Speech

Danwei Cai, Audrey Kittredge, Ben Naismith, Xiangying Jiang, Kevin Yancey


Abstract
This paper describes the development and validation of an automated scoring model for open-ended chatbot-based conversational speech among English learners in the Duolingo learning app. The model strongly predicted human ratings, produced reliable scores, and showed concurrent validity with Duolingo English Test speaking items, demonstrating stealth proficiency assessment at scale.
Anthology ID:
2026.aimecon-main.11
Volume:
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Month:
October
Year:
2026
Address:
Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States
Editors:
Joshua Wilson, Christopher Ormerod, Magdalen Beiting-Parrish
Venue:
AIME-Con
SIG:
Publisher:
National Council on Measurement in Education (NCME)
Note:
Pages:
100–108
Language:
URL:
https://aclanthology.org/2026.aimecon-main.11/
DOI:
Bibkey:
Cite (ACL):
Danwei Cai, Audrey Kittredge, Ben Naismith, Xiangying Jiang, and Kevin Yancey. 2026. Developing and Validating an Automatic Scoring Model for Chatbot-Based Conversational Speech. In Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers, pages 100–108, Wyndham Grand Pittsburgh Downtown, Pittsburgh, Pennsylvania, United States. National Council on Measurement in Education (NCME).
Cite (Informal):
Developing and Validating an Automatic Scoring Model for Chatbot-Based Conversational Speech (Cai et al., AIME-Con 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.aimecon-main.11.pdf