Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments

Oliver Cakebread-Andrews; Le An Ha; Ingo Frommholz; Burcu Can

Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments

Oliver Cakebread-Andrews, Le An Ha, Ingo Frommholz, Burcu Can

Abstract

This paper summarises the differences and similarities found between humans and three natural language processing models when attempting to identify whether English online comments are sarcastic or not. Three models were used to analyse 300 comments from the FigLang 2020 Reddit Dataset, with and without context. The same 300 comments were also given to 39 non-native speakers of English and the results were compared. The aim was to find whether there were any results that could be applied to English as a Foreign Language (EFL) teaching. The results showed that there were similarities between the models and non-native speakers, in particular the logistic regression model. They also highlighted weaknesses with both non-native speakers and the models in detecting sarcasm when the comments included political topics or were phrased as questions. This has potential implications for how the EFL teaching industry could implement the results of error analysis of NLP models in teaching practices.

Anthology ID:: 2024.lrec-main.552
Volume:: Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Month:: May
Year:: 2024
Address:: Torino, Italia
Editors:: Nicoletta Calzolari, Min-Yen Kan, Veronique Hoste, Alessandro Lenci, Sakriani Sakti, Nianwen Xue
Venues:: LREC | COLING
SIG:
Publisher:: ELRA and ICCL
Note:
Pages:: 6247–6256
Language:
URL:: https://aclanthology.org/2024.lrec-main.552/
DOI:
Bibkey:
Cite (ACL):: Oliver Cakebread-Andrews, Le An Ha, Ingo Frommholz, and Burcu Can. 2024. Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 6247–6256, Torino, Italia. ELRA and ICCL.
Cite (Informal):: Error Analysis of NLP Models and Non-Native Speakers of English Identifying Sarcasm in Reddit Comments (Cakebread-Andrews et al., LREC-COLING 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.lrec-main.552.pdf

PDF Cite Search Fix data