MaiChat: A Text-based Dialogue Corpus Rich in Conversational Features

Mai Hoang Dao, Catherine Lai, Peter Bell


Abstract
We present a new English corpus of typed instant-messaging dialogues that includes detailed timing information. Messages are collected from interactions between pairs who know each other well; the corpus is rich in typed features that augment the purely lexical, including hesitations, self-corrections, expressive respellings, and other markers of spontaneous interaction. Messages are collected using a custom-built chat platform that logs not only message content but also keystroke dynamics, screen activity, and demographic metadata. Designed with a transparent and reproducible protocol, the corpus enables scalable data collection while ensuring privacy and consent. We intend that the rich collection of features collected will facilitate future research in areas such as cognitive modelling, human–computer interaction, and conversational AI.
Anthology ID:
2026.lrec-1.123
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
1585–1594
Language:
External URL:
https://lrec.elra.info/lrec2026-main-123
DOI:
10.63317/3kpp3zj47d6d
Bibkey:
Cite (ACL):
Mai Hoang Dao, Catherine Lai, and Peter Bell. 2026. MaiChat: A Text-based Dialogue Corpus Rich in Conversational Features. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 1585–1594, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
MaiChat: A Text-based Dialogue Corpus Rich in Conversational Features (Dao et al., LREC 2026)
Copy Citation: