What Are the Odds? Language Models Are Capable of Probabilistic Reasoning

Akshay Paruchuri; Jake Garrison; Shun Liao; John Hernandez; Jacob Sunshine; Tim Althoff; Xin Liu; Daniel McDuff

What Are the Odds? Language Models Are Capable of Probabilistic Reasoning

Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, Daniel McDuff

Abstract

Language models (LM) are capable of remarkably complex linguistic tasks; however, numerical reasoning is an area in which they frequently struggle. An important but rarely evaluated form of reasoning is understanding probability distributions. In this paper, we focus on evaluating the probabilistic reasoning capabilities of LMs using idealized and real-world statistical distributions. We perform a systematic evaluation of state-of-the-art LMs on three tasks: estimating percentiles, drawing samples, and calculating probabilities. We evaluate three ways to provide context to LMs 1) anchoring examples from within a distribution or family of distributions, 2) real-world context, 3) summary statistics on which to base a Normal approximation. Models can make inferences about distributions, and can be further aided by the incorporation of real-world context, example shots and simplified assumptions, even if these assumptions are incorrect or misspecified. To conduct this work, we developed a comprehensive benchmark distribution dataset with associated question-answer pairs that we have released publicly.

Anthology ID:: 2024.emnlp-main.654
Volume:: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2024
Address:: Miami, Florida, USA
Editors:: Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 11712–11733
Language:
URL:: https://aclanthology.org/2024.emnlp-main.654
DOI:
Bibkey:
Cite (ACL):: Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, and Daniel McDuff. 2024. What Are the Odds? Language Models Are Capable of Probabilistic Reasoning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 11712–11733, Miami, Florida, USA. Association for Computational Linguistics.
Cite (Informal):: What Are the Odds? Language Models Are Capable of Probabilistic Reasoning (Paruchuri et al., EMNLP 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.emnlp-main.654.pdf
Software:: 2024.emnlp-main.654.software.zip
Data:: 2024.emnlp-main.654.data.zip

PDF Cite Search Software Data