Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process Generics

Emily Allaway; Chandra Bhagavatula; Jena D. Hwang; Kathleen McKeown; Sarah-Jane Leslie

doi:10.1162/coli_a_00530

Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process Generics

Emily Allaway, Chandra Bhagavatula, Jena D. Hwang, Kathleen McKeown, Sarah-Jane Leslie

Abstract

Large language models (LLMs) have garnered a great deal of attention for their exceptional generative performance on commonsense and reasoning tasks. In this work, we investigate LLMs’ capabilities for generalization using a particularly challenging type of statement: generics. Generics express generalizations (e.g., birds can fly) but do so without explicit quantification. They are notable because they generalize over their instantiations (e.g., sparrows can fly) yet hold true even in the presence of exceptions (e.g., penguins do not). For humans, these generic generalizations play a fundamental role in cognition, concept acquisition, and intuitive reasoning. We investigate how LLMs respond to and reason about generics. To this end, we first propose a framework grounded in pragmatics to automatically generate both exceptions and instantiations – collectively exemplars. We make use of focus—a pragmatic phenomenon that highlights meaning-bearing elements in a sentence—to capture the full range of interpretations of generics across different contexts of use. This allows us to derive precise logical definitions for exemplars and operationalize them to automatically generate exemplars from LLMs. Using our system, we generate a dataset of ∼370kexemplars across ∼17k generics and conduct a human validation of a sample of the generated data. We use our final generated dataset to investigate how LLMs reason about generics. Humans have a documented tendency to conflate universally quantified statements (e.g., all birds can fly) with generics. Therefore, we probe whether LLMs exhibit similar overgeneralization behavior in terms of quantification and in property inheritance. We find that LLMs do show evidence of overgeneralization, although they sometimes struggle to reason about exceptions. Furthermore, we find that LLMs may exhibit similar non-logical behavior to humans when considering property inheritance from generics.

Anthology ID:: 2024.cl-4.2
Volume:: Computational Linguistics, Volume 50, Issue 4 - December 2024
Month:: December
Year:: 2024
Address:: Cambridge, MA
Venue:: CL
SIG:
Publisher:: MIT Press
Note:
Pages:: 1211–1275
Language:
URL:: https://aclanthology.org/2024.cl-4.2/
DOI:: 10.1162/coli_a_00530
Bibkey:
Cite (ACL):: Emily Allaway, Chandra Bhagavatula, Jena D. Hwang, Kathleen McKeown, and Sarah-Jane Leslie. 2024. Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process Generics. Computational Linguistics, 50(3):1211–1275.
Cite (Informal):: Exceptions, Instantiations, and Overgeneralization: Insights into How Language Models Process Generics (Allaway et al., CL 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.cl-4.2.pdf

PDF Cite Search Fix data