The Linguist’s Lie Detector: Linguistic Knowledge in Large Language Models

Lucía Catalán Gris, Kim Gerdes, John S. Y. Lee


Abstract
We present a benchmark and evaluation pipeline for assessing how well large language models (LLMs) handle linguistic knowledge. Starting from a curated subcorpus of 11 syntax-focused articles published in Glossa: A Journal of General Linguistics (2016–2026), we design a pipeline that (1) segments article text into sentences, (2) extracts atomic, verifiable statements, and (3) classifies them into linguistic categories (language-specific, typological, theoretical, citation, or structural). Each stage is evaluated against human gold annotations produced by three annotators, with inter-annotator agreement measured via Krippendorff’s α and Cohen’s κ. We compare several LLMs on extraction and classification, using BERTScore-style similarity for extraction and macro F1 for classification. Finally, we generate contradictions of the true linguistic statements and test whether LLMs can distinguish true from false claims. On a challenge set of 705 linguistic statements, we compare eight LLMs, with Gemini 3 Flash achieving the highest F1 score of 0.66, indicating that current models possess limited but non-trivial linguistic knowledge.
Anthology ID:
2026.nslp-1.23
Volume:
Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Georg Rehm, Stefan Dietze, Danilo Dessi, Diana Maynard, Sonja Schimmler
Venues:
NSLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
235–246
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-nslp-23
DOI:
10.63317/2f7zobe6yo7h
Bibkey:
Cite (ACL):
Lucía Catalán Gris, Kim Gerdes, and John S. Y. Lee. 2026. The Linguist’s Lie Detector: Linguistic Knowledge in Large Language Models. In Proceedings of Natural Scientific Language Processing (NSLP) @ LREC 2026, pages 235–246, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
The Linguist’s Lie Detector: Linguistic Knowledge in Large Language Models (Catalán Gris et al., NSLP 2026)
Copy Citation: