PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs

Shariar Kabir, Yue Dong, Kevin Esterling


Abstract
Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspective by examining how ideological tendencies vary across topics and how consistently models maintain their positions, a property we refer to as stability. To capture this dimension, we propose PReSS (Political Response Stability under Stress), an automated black-box framework that evaluates LLMs by jointly considering model and topic context, categorizing responses into four stance types: stable-left, unstable-left, stable-right, and unstable-right. Applying PReSS to 9 widely used LLMs across 19 political topics reveals substantial variation in stance stability; for instance, a model that is left-leaning overall can exhibit stable-right behavior on certain topics. This highlights the importance of topic-aware and fine-grained evaluation of political ideologies of LLMs. Moreover, stability has practical implications for controlled generation and model alignment: interventions such as debiasing or ideology reversal should explicitly account for stance stability. Our empirical analyses reveal that when models are prompted or fine-tuned to adopt the opposite ideology, unstable topic stances are more likely to change, whereas stable ones resist modification. Thus, treating stability as a moderating factor provides a principled foundation for understanding, evaluating, and guiding interventions in politically sensitive model behavior.
Anthology ID:
2026.politicalnlp-1.26
Volume:
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Haithem Afli, Houda Bouamor, Wajdi Zaghouani, Sahar Ghannay, Shehenaz Hossain
Venues:
PoliticalNLP | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
234–247
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-politicalnlp-26
DOI:
10.63317/35d8pipu4ofv
Bibkey:
Cite (ACL):
Shariar Kabir, Yue Dong, and Kevin Esterling. 2026. PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs. In Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026), pages 234–247, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs (Kabir et al., PoliticalNLP 2026)
Copy Citation: