Natural Language Processing for Political Sciences (2026)


up

pdf (full)
bib (full)
Proceedings of the 3rd Workshop on Natural Language Processing for Political Sciences (PoliticalNLP 2026)

Civil-society organisations, journalists, and fact-checkers monitoring elections require scalable ways to convert high-volume political news into actionable narrative intelligence, yet most NLP pipelines stop at classification outputs that are difficult to operationalise. We present a methodology-driven case study assessing whether large language models, constrained by an explicit analytical schema and multi-stage validation, can reliably transform Hungarian pre-election news into structured narrative intelligence briefs. Using RSS-scraped content from 21 Hungarian-language sources (574 election-relevant articles), we implement a multi-stage pipeline: (1) per-article extraction of narrative event frames grounded in the Narrative Policy Framework (actor–action–target with role assignment and causal claims) and manipulation techniques from the SemEval propaganda taxonomy; (2) embedding-based clustering of narrative frames with domain classification; and (3) constrained brief generation producing five structured sections—narrative summary, character map, manipulation profile, escalation assessment, and counter-strategy—where counter-strategies are grounded in verified external sources via curated contextual cards and constrained by evidence-based de-escalation principles. We evaluate brief quality through dual-track evaluation combining three human domain experts and three LLM judges on a single brief, with a scaled 29-brief LLM-as-judge assessment, and document key failure modes across a defined taxonomy. We conclude with implications for trustworthy human-in-the-loop political NLP and the practical limits of LLM-assisted narrative intelligence.
The growing complexity and diversity of news coverage have made framing analysis a crucial yet challenging task in computational social science. Traditional approaches, including manual annotation and fine-tuned models, remain limited by high annotation costs, domain specificity, and inconsistent generalisation. Instruction-based large language models (LLMs) offer a promising alternative, yet their reliability for framing analysis remains insufficiently understood. In this paper, we conduct a systematic evaluation of several LLMs, including GPT-3.5/4, FLAN-T5, and Llama 3, across zero-shot, few-shot, and explanation-based prompting settings. Focusing on domain shift and inherent annotation ambiguity, we show that model performance is highly sensitive to prompt design and prone to systematic errors on ambiguous cases. Although LLMs, particularly GPT-4, exhibit stronger cross-domain generalisation, they also display systematic biases, most notably a tendency to conflate emotional language with framing. To enable principled evaluation under real-world topic diversity, we introduce a new dataset of out-of-domain news headlines covering diverse subjects. Finally, by analysing agreement patterns across multiple models on existing framing datasets, we demonstrate that cross-model consensus provides a useful signal for identifying contested annotations, offering a practical approach to dataset auditing in low-resource settings.
African Twitter users are active shapers of the Palestine-Israel conversation but their contribution remains relatively understudied. Using 132.5K geo-located tweets from 2020 to 2023 and 451-term list of keywords in 33 languages, we identify three patterns in this context: (1) broad participation (Egypt supplies 43% of posts, yet Nigeria, South Africa, Kenya and Ghana contribute more than a third); (2) multilingual predominantly pro-Palestine amplification across Arabic, English, French, Swahili and other tongues, with 8% of tweets left “undetermined” by Twitter’s language detector; and (3) a humanitarian framing that centers civilian harm through hashtags such as #GazaUnderAttack and #PalestenianLivesMatter. We outline design implications for language-agnostic interfaces, low-friction source verification and cross-movement recommendation tools that foreground African epistemologies in global civic-tech systems.
YouTube Shorts have become central to news consumption on the platform, yet research on how geopolitical events are represented in this format remains limited. To address this gap, we present a multimodal pipeline that combines automatic transcription, aspect-based sentiment analysis (ABSA), and semantic scene classification. The pipeline is first assessed for feasibility and then applied to analyze short-form coverage of the Israel–Hamas war by state-funded outlets. Using over 2,300 conflict-related Shorts and more than 94,000 visual frames, we systematically examine war reporting across major international broadcasters. Our findings reveal that the sentiment expressed in transcripts regarding specific aspects differs across outlets and over time, whereas scene-type classifications reflect visual cues consistent with real-world events. Notably, smaller domain-adapted models outperform large transformers and even LLMs for sentiment analysis, underscoring the value of resource-efficient approaches for humanities research. The pipeline serves as a template for other short-form platforms, such as TikTok and Instagram, and demonstrates how multimodal methods, combined with qualitative interpretation, can characterize sentiment patterns and visual cues in algorithmically driven video environments.
Online news platforms have become central spaces for political discourse, playing a critical role in shaping public opinion and democratic participation. In large-scale democracies such as India, the combination of extensive user engagement, ideological polarization, and automated participation raises significant concerns regarding trust, transparency, and manipulation. This paper presents a large-scale empirical study of political discourse on a prominent Indian news platform, analyzing over 21,000 news articles and more than 1.5 million user comments. We investigate how ideological bias, sentiment dynamics, and non-organic user behavior interact to shape engagement patterns. Our methodology integrates hybrid article bias classification, large-scale sentiment analysis, heuristic-based bot detection, coordinated behavior analysis, and a focused examination of super-active users. In addition, we compare rule-based stance inference with large language model (LLM)-based stance classification to assess trade-offs between computational efficiency and contextual accuracy. The results reveal systematic sentiment skew, disproportionate influence by super-active and bot-like users, and coordinated campaigns aligned with specific political narratives. We conclude by discussing the implications of these findings for trust in online political discourse and reflecting on the dual role of generative AI as both an analytical tool and a potential vector for manipulation.
Social media platforms play a pivotal role in shaping public opinion and amplifying political discourse, particularly during elections. However, the same dynamics that foster democratic engagement can also exacerbate polarization. To better understand these challenges, here, we investigate the ideological positioning of tweets related to the 2024 U.S. Presidential Election. To this end, we analyze 1,235 tweets from key political figures and 63,322 replies, and classify ideological stances into Pro-Democrat, Anti-Republican, Pro-Republican, Anti-Democrat, and Neutral categories. Using a classification pipeline involving three large language models (LLMs)—GPT-4o, Gemini-Pro, and Claude-Opus—and validated by human annotators, we explore how ideological alignment varies between candidates and constituents. We find that Republican candidates author significantly more tweets in criticism of the Democratic party and its candidates than vice versa, but this relationship does not hold for replies to candidate tweets. Furthermore, we highlight shifts in public discourse observed during key political events. By shedding light on the ideological dynamics of online political interactions, these results provide insights for policymakers and platforms seeking to address polarization and foster healthier political dialogue.
We present a unified pipeline for synthesizing high-quality Quechua and Spanish speech for the Peruvian Constitution using three state-of-the-art text-to-speech (TTS) architectures: XTTS v2, F5-TTS, and DiFlow-TTS. Our models are trained on independent Spanish and Quechua speech datasets with heterogeneous sizes and recording conditions, and leverage bilingual and multilingual TTS capabilities to improve synthesis quality in both languages. By exploiting cross-lingual transfer, our framework mitigates data scarcity in Quechua while preserving naturalness in Spanish. We release trained checkpoints, inference code, and synthesized audio for each constitutional article, providing a reusable resource for speech technologies in indigenous and multilingual contexts. This work contributes to the development of inclusive TTS systems for political and legal content in low-resource settings
Parliamentary debate constitutes a central arena of political power, shaping legislative outcomes and public discourse. Incivility within this arena signals political polarization and institutional conflict. This study presents a systematic investigation of incivility in the German Bundestag by examining calls to order (CtO; plural: CtOs) as formal indicators of norm violations. Despite their relevance, CtOs have received little systematic attention in parliamentary research. We introduce a rule-based method for detecting and annotating CtOs in parliamentary speeches and present a novel dataset of German parliamentary debates spanning 72 years that includes annotated CtO instances. Additionally, we develop the first classification system for CtO triggers and analyze the factors associated with their occurrence. Our findings show that, despite formal regulations, the issuance of CtOs is partly subjective and influenced by session presidents and parliamentary dynamics, with certain individuals disproportionately affected. An insult towards individuals is the most frequent cause of CtO. In general, male members and those belonging to opposition parties receive more calls to order than their female and coalition-party counterparts. Most CtO triggers were detected in speeches dedicated to governmental affairs and actions of the presidency. The CtO triggers dataset is available at: https://github.com/kalawinka/cto_analysis.
Large Language Models (LLMs) are increasingly deployed in politically sensitive environments, where memorisation of personal data or confidential content raises regulatory concerns under frameworks such as the GDPR and its “right to be forgotten”. Translating such legal principles into large-scale generative systems presents significant technical challenges. We introduce a lightweight sequential unlearning framework that explicitly separates retention and suppression objectives. The method first stabilises benign capabilities through positive fine-tuning, then applies layer-restricted negative fine-tuning to suppress designated sensitive patterns while preserving general language competence. Experiments on the SemEval-2025 LLM Unlearning benchmark demonstrate effective behavioural suppression with minimal impact on factual accuracy and fluency. GPT-2 exhibits greater robustness than DistilGPT-2, highlighting the role of model capacity in privacy-aligned adaptation. We position sequential unlearning as a practical and reproducible mechanism for operationalising data erasure requirements in politically deployed LLMs.
Political scientists are interested in changes in political discourse over time. However, the topics of interest, such as the changes in support for or understanding of certain narratives, are often ill-defined and require deliberation, which prevents most lexical or metadata-based methods of temporal aggregation. To enable a diachronic analysis, we propose to model such settings as a series of binary document classification tasks – which current reasoning LLMs can adequately solve – and aggregate the decisions into a temporal signal. Specifically, we propose to use LLMs to classify if a parliament speech is in support of either of two narratives, and we use the monthly count of positives per narrative to track the support over time. We show that the classification is sufficiently accurate and use it to create detailed time series data showing support for the selected narratives in speeches given in the European Parliament from 2006 to 2023. The method is developed in close collaboration with political scientists and is considered an ideal starting point for diachronic analyses of political decision-making processes by domain experts.
This paper introduces ParlaCAP, a large-scale dataset for analyzing parliamentary agenda setting across Europe, and proposes a cost-effective method for building domain-specific policy topic classifiers. Applying the Comparative Agendas Project (CAP) schema to the multilingual ParlaMint corpus of over 8 million speeches from 28 parliaments of European countries and autonomous regions, we follow a teacher-student framework in which a high-performing large language model (LLM) annotates in-domain training data and a multilingual encoder model is fine-tuned on these annotations for scalable data annotation. We show that this approach produces a classifier tailored to the target domain. Agreement between the LLM and human annotators is comparable to inter-annotator agreement among humans, and the resulting model outperforms existing CAP classifiers trained on manually-annotated but out-of-domain data. In addition to the CAP annotations, the ParlaCAP dataset offers rich speaker and party metadata, as well as sentiment predictions coming from the ParlaSent multilingual transformer model, enabling comparative research on political attention and representation across countries. We illustrate the analytical potential of the dataset with three use cases, examining the distribution of parliamentary attention across policy topics, sentiment patterns in parliamentary speech, and gender differences in policy attention.
Understanding how political orientation influences lexical choices is essential for detecting bias and framing in news media. In this paper, we present a computational framework for identifying nouns whose interpretation varies across politically divergent newspapers. Using a large corpus of French news articles published in 2024, we categorize texts by topics and political orientation. We use contextual embeddings to cluster occurrences of nouns to detect semantic variations and dissimilarity among sources. This allows us to map semantic distances between newspapers and identify polarized or editorially marked lexical choices. Our results show that topics, polysemy, and editorial priorities contribute differently to lexical divergence. We discuss these findings and highlight how contextual embeddings can help reveal semantic biases that would remain invisible through frequency-based methods. We conclude by outlining perspectives for improving topic classification and the clustering method, exploring alternative divergence measures, conducting a qualitative analysis of our results, and extending the framework to other languages or genres.
Policy issues are central to election campaigns, yet systematic analyses of issue communication on Instagram remain scarce — particularly for ephemeral Stories. We develop and evaluate automated methods for detecting the binary presence of policy issues in Instagram posts and Stories from the 2021 German federal election. Drawing on a gold-standard dataset of 1,357 annotated documents across three textual channels (captions, OCR-extracted image text, and speech transcripts), we compare a fine-tuned German transformer (GBERT) with multiple LLM prompting strategies (zero-shot, few-shot, retrieval-augmented). Both approaches prove effective: GBERT achieves a cross-validated macro F1 of 0.90, closely matched by GPT-o3 under few-shot prompting (0.88). Substantively, policy visibility varies far more by content format than by party: 70% of posts contain policy references compared to only 17% of Stories, a pattern that holds consistently across all eight parties. An exploratory topic model confirms that parties reproduce familiar issue-ownership profiles within the subset of policy-relevant texts. Our results establish binary issue detection as a feasible foundation for studying policy communication in multimodal, ephemeral social media environments.
Calls-to-action (CTAs) are central to digital campaigning, yet computational research has largely focused on binary detection only. We address CTA type classification in German Instagram campaign texts (posts and ephemeral stories), distinguishing Support, Inform, Interact, and No CTA. With limited annotated data, we benchmark a fine-tuned GBERT model against GPT models using zero-shot, few-shot, and retrieval-augmented few-shot prompting in a multi-label setup. Both approaches reach similar performance in five-fold cross-validation (macro-F1 ca. 0.79), with persistent difficulty on the rare Interact category. As a proof of concept, we apply the selected setup to the 2021 federal election corpus and show that parties varied not only in overall CTA use but also in how they balanced appeals across posts versus stories. The results demonstrate the feasibility of CTA type classification with modest data and position retrieval-augmented prompting as a practical alternative to supervised fine-tuning.
This paper investigates the correlation of hate speech and hate crime, in a inter-disciplinary approach, using a computational hate speech and hate crime detection method in a socio-political science framework, coupling Natural Language Processing with Political Sciences. The study focuses on Greece in the turbulent period from 2015 to 2022 (a period marked by economic, refugee, foreign policy, and pandemic crises); it analyzes tweets to discern linguistic patterns used to verbally attack predefined target groups consisting of ethnic and religious minorities in the country. Furthermore, it investigates hate crimes reported in the press, against the same target groups and during the same period and proceeds to examine correlations between xenophobic attitudes expressed verbally through social media, and those manifested as physical attacks in real life.
Navigating AI regulation across jurisdictions is increasingly difficult for policymakers, legal professionals, and researchers. To address this, we present a multi-jurisdictional Retrieval-Augmented Generation system for global AI regulation. Our corpus includes 241 documents across 73 jurisdictions, ranging from formal legislation like the EU AI Act to unstructured policy documents such as national AI strategies. The system makes three technical contributions: type-specific chunking that preserve legal structure across heterogenous documents; conditional retrieval routing with entity detection and metadata for legal citations; and priority-based re-ranking to boost enacted legislation over policy and secondary sources. Evaluation of 50 queries reveals strong performance across both single-entity and multi-jurisdictional questions, achieving 0.87 average faithfulness and 0.84 average answer relevancy. Single-entity queries achieve 0.86 average faithfulness and 0.92 average answer relevancy, while multi-jurisdictional comparison queries achieve 0.88 average faithfulness and 0.75 average answer relevancy. These findings highlight the effectiveness of domain-specific retrieval strategies for navigating complex, heterogenous regulatory corpora.
Sentiment analysis remains the dominant computational approach for evaluating political news, yet its ability to capture the rhetorical complexity of political discourse is increasingly questioned. This paper presents a systematic comparison between a transformer-based sentiment classifier (RoBERTa) and a Large Language Model-based multi-dimensional framing analysis framework across 50 political news articles from 17 international outlets. While RoBERTa classifies 70% of articles as neutral and reduces political discourse to a three-way polarity scale, the LLM-based framework captures 13 numerical dimensions including bias direction and intensity, manipulation indicators (cherry-picking, loaded language, false equivalence), sensationalism, and communicative intent. Our correlation analysis reveals only weak-to-moderate relationships between sentiment polarity and framing dimensions (maximum Pearson r = 0.38, p < 0.01), demonstrating that these approaches measure fundamentally different properties of political text. Through case studies, we show that sentiment-neutral articles can exhibit extreme manipulation patterns, while highly negative articles may reflect factual reporting on inherently negative events. These findings argue for moving beyond sentiment as a proxy for media quality, toward multi-dimensional frameworks that can reveal the rhetorical strategies invisible to polarity-based analysis. All data and analysis code will be made available upon acceptance.
This research investigates the pragmatic competence of Large Language Models (LLMs) in interpreting implicit meanings within Italian political discourse. Using the IMPAQTS-PIDMM dataset, which is a multimodal benchmark derived from the 2.5-million-token IMPAQTS corpus, the experiment evaluates how effectively models identify tendentious content such as presuppositions and implicatures. The study compares the performance of text-only LLMs against speech-based models (SpeechLMs) that process both audio and transcriptions to determine if acoustic cues enhance understanding. The results reveal that text-only models significantly outperform multimodal variants, with Qwen2.5-72B achieving the highest global accuracy of 0.863. Surprisingly, the inclusion of audio did not improve performance, as SpeechLMs like GPT-4o-mini-audio-preview and Qwen2-Audio-7B-Instruct obtained lower accuracy scores and a higher frequency of missed answers compared to their text-only equivalents. Across all tested architectures, models generally demonstrated a superior ability to process presuppositions over implicatures.
The integration of artificial intelligence (AI), particularly large language models (LLMs), into research across the social sciences has accelerated innovation but also introduced significant challenges to reproducibility - a cornerstone of scientific integrity. In this review of scientific practices, we examine the reproducibility crisis in AI-driven research with a focus on psychology, identifying common pitfalls, reviewing proposed solutions, and advocating for best practices. Common pitfalls in current practices in the social sciences are identified and highlighted through synthesized research scenarios, such as: (1) using inaccessible datasets or language models with restricted access, (2) treating black-box API outputs as stable observations ignoring updates and hidden changes, (3) producing single runs for measurements instead of stochastic draws for aggregated performances, (4) failing to report full LLM version, prompting, and sampling parameters, and (5) opaque training and fine-tuning of LLMs. Our recommended practices include precisely documenting the model used, fixing all inference parameters, using automation and scripts to control prompts, context, and outputs, and standardizing the environment and API conditions. By embracing transparency and methodological rigour, we can transform the challenges of AI-driven research into opportunities for more robust and impactful science, ensuring that innovation never comes at the cost of credibility.
This study analyzes a publicly released dataset from a discontinued field experiment on Reddit’s r/ChangeMyView. The intervention, conducted by unknown, external researchers and halted following ethical backlash, involved undisclosed AI-generated accounts engaging users in live debate. After public disclosure, Reddit authorized moderators to release an archive of the AI-generated comments, creating a rare opportunity to examine how large language models operated in an identity-rich deliberative forum without disclosure. We conduct a structured content analysis of this corpus, evaluating identity performance, authority signaling, alignment strategies, and activation of cognitive heuristics. Identity targeting or adoption appears in over two-thirds of comments, alignment moves and authority claims in nearly all of them, and cognitive-bias triggers—particularly confirmation bias, representativeness, and availability—in the large majority. These patterns co-occur systematically, composing a rhetorical architecture calibrated for persuasive efficiency rather than authentic deliberative participation. Compared against human-authored CMV counter-arguments, the agents inverted the typical distribution on every dimension: denser authority use, more adversarial alignment, and heavier reliance on external citation over experiential grounding. In such environments, distinctions between authentic and synthetic epistemic standing grow increasingly opaque—an asymmetry that disclosure mandates alone cannot address. The results point toward auditing frameworks capable of assessing how AI systems structure credibility, not merely whether they are present. Our dataset is available at https://github.com/kokiljaidka/UnauthorizedRedditCMVPosts
This paper examines the party leanings of international and Norwegian large language models. The experiments are two-fold; first they are asked to answer the question of a Valgomat–an election affiliation guide–as a neutral observer, and secondly as if it were a paying party member of the parties in the data. Results show that the neutral prompting show centrist leanings, whereas models struggle with mimicking party members. Models with additional training on Norwegian text performed better on this task.
We investigate automatic attitude detection in UN Security Council speeches using adapters. Following Martin and White’s Appraisal Theory, we identify three types of evaluative language: affect (emotional responses such as hope or concern), judgement (ethical evaluations of behavior), and appreciation (valuations of objects or situations). Training only 0.95% of BERT-large’s parameters, adapters achieve F1 scores ranging from 0.76 (affect) to 0.46 (appreciation), approaching full fine-tuning performance while enabling rapid task-specific experimentation. Differences in observed evaluation metrics mirror the pattern of the human inter-annotator agreement. This correlation suggests that computational difficulty reflects genuine linguistic ambiguity. Affect benefits from conventionalized diplomatic expressions, while appreciation faces context-dependent evaluation and severe class imbalance. Analysis demonstrates that evaluative intensity varies systematically across diplomatic contexts, with implications for corpus design in specialized discourse analysis.
As large language models (LLMs) increasingly mediate political information across linguistic contexts, concerns emerge regarding cross-lingual consistency in political framing. We investigate whether multilingual LLMs generate systematically different rhetorical and semantic frames when prompted in English versus Arabic on the politically salient issue of migration. Focusing on two widely used models, i.e., GPT-4o and Jais-13B, we implement a controlled prompt design (N = 800 generations; 400 per language), to isolate language as the primary experimental variable. We introduce a mixed method evaluation framework that combines lexical frame analysis, statistical association testing, and qualitative discourse analysis. Our results show a significant association between language and framing distribution (χ2 = 43.32, p = 2.11 × 10−9). While security-oriented framing is prominent in both languages, English generations exhibit substantially higher rates of institutional and legislative framing, whereas Arabic generations show greater concentration in security and communitarian discourse. These findings indicate that input language acts as a conditioning signal that systematically modulates political framing within multilingual LLMs, even under controlled semantic prompts. We conceptualize this phenomenon as cross-lingual framing drift and discuss its implications for multilingual alignment, political bias evaluation, and global information ecosystems. We conclude by outlining an evaluative protocol for detecting language-conditioned asymmetries in generative models. We make all data, code, and experimental settings publicly available at: https://github.com/NRAwwad/-A-Cross-Lingual-Analysis-of-Political-Framing-in-English-and-Arabic.git.
Sentiment analysis models are increasingly deployed to analyze political discourse, yet strong in-domain performance does not guarantee robustness under domain shift. We study cross-domain generalization in Hinglish (Hindi–English code-mixed) sentiment analysis by evaluating a fine-tuned XLM-RoBERTa classifier, trained on 29,000 general-domain Hinglish sentences, on a curated benchmark of politically oriented Hinglish text. While the model achieves 92.02% accuracy in-domain, performance drops to 71.83% under political domain shift. Error analysis reveals a pronounced directional bias with 48.9% of neutral political statements misclassified as negative, indicating a systematic neutrality-to-negative shift. In addition, 87.5% of incorrect predictions are assigned confidence scores above 95%, pointing to severe miscalibration under distribution shift. We further compare these results against an instruction-tuned large language model (Llama 3.3), which achieves 90.85% zero-shot accuracy and 94.37% accuracy with contextual prompting, while substantially reducing neutrality bias. Our findings indicate the need for domain-aware evaluation, calibration diagnostics, and explicit reporting of failure modes when deploying sentiment models in politically sensitive settings.
A comprehensive framework was developed to detect political bias in Arabic news articles, with a case study focusing on media reporting of the Palestinian issue. The methodology integrates MARBERT contextual embeddings with classical and deep learning classifiers, including SVM, Logistic Regression, Random Forest, and LSTM. The scalability of data processing was ensured through Apache Spark for potential real-time deployment. Experimental results showed that fine-tuned MARBERT embeddings combined with LSTM achieved the highest classification accuracy of 0.87, along with notable improvements in F1-scores across the pro, against, and neutral categories. These findings highlight the effectiveness of domain-specific fine-tuning of transformer models for political bias classification. The study also addressed class imbalance using SMOTE and class weighting strategies, and assessed feature robustness using multiple vectorization techniques.
Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspective by examining how ideological tendencies vary across topics and how consistently models maintain their positions, a property we refer to as stability. To capture this dimension, we propose PReSS (Political Response Stability under Stress), an automated black-box framework that evaluates LLMs by jointly considering model and topic context, categorizing responses into four stance types: stable-left, unstable-left, stable-right, and unstable-right. Applying PReSS to 9 widely used LLMs across 19 political topics reveals substantial variation in stance stability; for instance, a model that is left-leaning overall can exhibit stable-right behavior on certain topics. This highlights the importance of topic-aware and fine-grained evaluation of political ideologies of LLMs. Moreover, stability has practical implications for controlled generation and model alignment: interventions such as debiasing or ideology reversal should explicitly account for stance stability. Our empirical analyses reveal that when models are prompted or fine-tuned to adopt the opposite ideology, unstable topic stances are more likely to change, whereas stable ones resist modification. Thus, treating stability as a moderating factor provides a principled foundation for understanding, evaluating, and guiding interventions in politically sensitive model behavior.
Political polarization emerges from a complex interplay of beliefs about policies, figures, and issues. However, most computational analyses reduce discourse to coarse partisan labels, overlooking how these beliefs interact. This is especially evident in online political conversations, which are often nuanced and cover a wide range of subjects, making it difficult to automatically identify the target of discussion and the opinion expressed toward them. In this study, we investigate whether Large Language Models (LLMs) can address this challenge through Target-Stance Extraction (TSE), a recent natural language processing task that combines target identification and stance detection, enabling more granular analysis of political opinions. For this, we construct a dataset of 1,084 Reddit posts from r/NeutralPolitics, covering 138 distinct political targets and evaluate a range of proprietary and open-source LLMs using zero-shot, few-shot, and context-augmented prompting strategies. Our results show that the best models perform comparably to highly trained human annotators and remain robust on challenging posts with low inter-annotator agreement. These findings demonstrate that LLMs can extract complex political opinions with minimal supervision, offering a scalable tool for computational social science and political text analysis.
This study evaluates whether Large Language Models (LLMs) can facilitate qualitative political narrative analysis by comparing outputs from four models—Mistral, Llama, ChatGPT-4o, and DeepSeek—against narrative analyses written by expert scholars. Using European Union State of the Union speeches (2010–2023), we examine migration and solidarity narratives through semantic and lexical similarity metrics alongside systematic validation. The narrative scholars demonstrate strong semantic alignment despite differences in wording, establishing a benchmark for interpretive consistency. Across both topics, the models produce lexical and semantic similarity scores that are broadly comparable to those observed between the scholars themselves, with differences at these levels often marginal. However, similarity metrics do not provide the full picture. Validation reveals model-specific weaknesses that are not captured by lexical or semantic alignment alone, including factual errors, over-structural abstraction, and difficulty engaging less salient narrative threads. These findings demonstrate that LLMs can produce narratives that align closely with human outputs in semantic and lexical similarity, yet these measures alone are insufficient to assess interpretive quality.
The integration of large language models into political discourse analysis creates new opportunities for comparative research, policy analysis, and civic technology, while introducing material risks for democratic accountability. This paper argues that cultural adaptation is a prerequisite for trustworthy deployment of large language models in political communication across diverse linguistic and institutional contexts. Current systems remain shaped by English dominant data, uneven multilingual coverage, and assumptions grounded in a narrow range of political institutions and discourse conventions, producing systematic errors when applied across cultures. We formalize cultural adaptation across translation, discourse, and ontology levels, identify recurring cultural failure modes in political NLP, and propose an operational evaluation matrix grounded in cultural fidelity, calibration, and democratic safety. Building on political text analysis, sociotechnical auditing, and cross cultural pragmatics, we outline methodological pathways including participatory dataset development, culturally aware transfer learning, and benchmark design that makes cultural adaptation empirically measurable. We conclude by clarifying governance constraints and scope conditions under which culturally adaptive political NLP can support democratic legitimacy.
We present an automated crosswalk framework that compares an AI safety policy document pair under a shared taxonomy of activities. Using the activity categories defined in Activity Map on AI Safety as fixed aspects, the system extracts and maps relevant activities, then produces for each aspect a short summary for each document, a brief comparison, and a similarity score. We assess the stability and validity of LLM-based crosswalk analysis across public policy documents. Using five large language models, we perform crosswalks on ten publicly available documents and visualize mean similarity scores with a heatmap. The results show that model choice substantially affects the crosswalk outcomes, and that some document pairs yield high disagreements across models. A human evaluation by three experts on two document pairs shows high inter-annotator agreement, while model scores still differ from human judgments. These findings support comparative inspection of policy documents.