CausalGym: Benchmarking causal interpretability methods on linguistic tasks

Aryaman Arora; Dan Jurafsky; Christopher Potts

CausalGym: Benchmarking causal interpretability methods on linguistic tasks

Aryaman Arora, Dan Jurafsky, Christopher Potts

Abstract

Language models (LMs) have proven to be powerful tools for psycholinguistic research, but most prior work has focused on purely behavioural measures (e.g., surprisal comparisons). At the same time, research in model interpretability has begun to illuminate the abstract causal mechanisms shaping LM behavior. To help bring these strands of research closer together, we introduce CausalGym. We adapt and expand the SyntaxGym suite of tasks to benchmark the ability of interpretability methods to causally affect model behaviour. To illustrate how CausalGym can be used, we study the pythia models (14M–6.9B) and assess the causal efficacy of a wide range of interpretability methods, including linear probing and distributed alignment search (DAS). We find that DAS outperforms the other methods, and so we use it to study the learning trajectory of two difficult linguistic phenomena in pythia-1b: negative polarity item licensing and filler–gap dependencies. Our analysis shows that the mechanism implementing both of these tasks is learned in discrete stages, not gradually.

Anthology ID:: 2024.acl-long.785
Volume:: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: August
Year:: 2024
Address:: Bangkok, Thailand
Editors:: Lun-Wei Ku, Andre Martins, Vivek Srikumar
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 14638–14663
Language:
URL:: https://aclanthology.org/2024.acl-long.785
DOI:
Bibkey:
Cite (ACL):: Aryaman Arora, Dan Jurafsky, and Christopher Potts. 2024. CausalGym: Benchmarking causal interpretability methods on linguistic tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14638–14663, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):: CausalGym: Benchmarking causal interpretability methods on linguistic tasks (Arora et al., ACL 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.acl-long.785.pdf

PDF Cite Search