Workshop on Social Context (2026)
up
Proceedings of the 1st Workshop on Social Context (SoCon) and the 2nd Workshop on Integrating NLP and Psychology to Study Social Interactions (NLPSI) @ LREC 2026
Proceedings of the 1st Workshop on Social Context (SoCon) and the 2nd Workshop on Integrating NLP and Psychology to Study Social Interactions (NLPSI) @ LREC 2026
Marco Antonio Stranisci | Neele Falk | Sofie Labat | Soda Marem Lo | Aswathy Velutharambath | Sabine Weber | Rossana Damiano | Simona Frenda | Veronique Hoste | Bennett Kleinberg | Roman Klinger | Viviana Patti | Flor Miriam Plaza-del-Arco | Maarten Sap | Seid Muhie Yimam
Marco Antonio Stranisci | Neele Falk | Sofie Labat | Soda Marem Lo | Aswathy Velutharambath | Sabine Weber | Rossana Damiano | Simona Frenda | Veronique Hoste | Bennett Kleinberg | Roman Klinger | Viviana Patti | Flor Miriam Plaza-del-Arco | Maarten Sap | Seid Muhie Yimam
State vs. Trait Anxiety in Causal Language Models
Karin Shistik | Idan-Chaim Cohen | Aviad Elyashar | Ortal Slobodin | Odeya Cohen | Rami Puzis
Karin Shistik | Idan-Chaim Cohen | Aviad Elyashar | Ortal Slobodin | Odeya Cohen | Rami Puzis
Psychological constructs in humans range along a state–trait continuum: traits persist across situations, while states fluctuate with context. Studies have shown that language models exhibit measurable psychological constructs, yet whether these constructs differ in contextual stability, as the state–trait distinction predicts, remains untested. We present the Questionnaire for Causal Language Models (QCLM), a psychometric framework that measures constructs through next-token probability distributions of base models. Applying QCLM to 35 causal language models under vanilla, stress, and neutral conditions, we assess two anxiety instruments targeting opposite ends of the state–trait continuum: STAI-S (state anxiety) and STAI-T (trait anxiety). Paired effect sizes and variance decomposition reveal that state anxiety is more sensitive to stress manipulation than trait anxiety: stimulus type accounts for a larger share of variance in state anxiety, while model identity contributes more to trait anxiety. These results provide empirical evidence that the state–trait distinction extends to language model behavior.
Documenting Rural Gatherings in Aging Japan: Social Context and Language Use in Interaction at a Mobile Supermarket
Haruka Sakai | Rui Sakaida
Haruka Sakai | Rui Sakaida
This paper presents a documentation framework and an exploratory analysis of language use in everyday interactions at rural gatherings in aging Japan, a communicative setting shaped by distinct social contexts that remain largely absent from existing language resources. Drawing on studies of face-to-face encounters, we propose a typology of rural gatherings and examine mobile supermarkets (vehicles that transport and sell daily necessities at scheduled stops in areas that lack fixed retail stores) as a case study. We present a preliminary analysis based on a community-mediated recording methodology. The quantitative findings reveal that conversational hot spots occur immediately following the encounter and transaction phases, indicating that participants experience these encounters as occasions for social connection rather than mere commercial transactions. The qualitative findings from the interaction analysis demonstrate how participants simultaneously manage work and conversation through vocal, bodily, and temporal resources in a social context. We discuss how these findings illuminate dimensions in social contexts that require interdisciplinary investigation beyond what existing language resources currently capture.
This paper argues that, when it comes to modeling language variation and change on the video game streaming platform Twitch, it is necessary to consider “meso-level” communities of practice, i.e. communities of practice that are smaller than the full video game community, yet larger than the usual level of analysis in recent linguistics studies: communities associated with individual Twitch channels. We present a computational method for identifying these linguistically relevant communities of practice and show how this method can be useful for analyzing quantitative patterns of sociolinguistic variation in a corpus composed of the chat transcripts of 15 streamers of the game Elden Ring: Nightreign.
Implicit Cultural Identity Signals in Language: Detection and Effects in Negotiation Dialogue
Bin Han | Danah Yun | James Hale | Jonathan Gratch
Bin Han | Danah Yun | James Hale | Jonathan Gratch
Language conveys cultural identity even when not intentionally disclosed. This study examines cultural signals in task-oriented dialogue using English negotiations from the KODIS dataset. We focus our analysis on participants from four countries: the US, UK, Mexico, and South Korea. Interacting anonymously under identical conditions, we evaluated whether a speaker’s country could be inferred from dialogue by zero-shot LLMs and embedding-based classifiers. Results show that while objective negotiation outcomes remained similar across groups, subjective perceptions varied significantly. Embedding-based models reliably identified country of origin, whereas zero-shot LLM performance dropped under distribution shift. These findings suggest that cultural identity-related signals are embedded in language and may be relevant for analyzing negotiation dialogue.
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
Tatiana Petrova | Stanislav Sokol | Radu State
Tatiana Petrova | Stanislav Sokol | Radu State
Which persuasion strategies, if any, are associated with donation compliance? Answering this requires fine-grained strategy labels across a full corpus and statistical tests corrected for multiple comparisons. We annotate all 10,600 persuader turns in the 1,017-dialogue PersuasionForGood corpus with a taxonomy of 41 strategies in 11 categories, using three open-source large language models (Qwen3:30b, Mistral-Small-3.2, Phi-4). Strategy categories alone explain little variance in donation outcome (pseudo R-squared approximately 0.015, consistent across all three annotators). Guilt Induction is the only strategy significantly associated with lower donation rates (approximately -23 percentage points), an effect that replicates across all three models despite only moderate inter-model agreement. Reciprocity is the most robust positive correlate. Target sentiment and interest predict whether a donation occurs but show at most a weak correlation with donation amount. Logistic regression with sentiment, interest, Guilt Induction, and Reciprocity achieves nearly the same fit (pseudo R-squared = 0.080) as the full model with all strategy categories. These findings suggest that strategy identification alone is insufficient to explain persuasion effectiveness, and that guilt-based appeals may be counterproductive in prosocial settings. We release the fully annotated corpus as a public resource.
Do LLMs Ask the Right Questions? Evaluating GPT-Generated Surveys as Instruments for Measuring Social Attitudes
Tina Behzad | Wenbo Li | Reuben Kline | Klaus Mueller
Tina Behzad | Wenbo Li | Reuben Kline | Klaus Mueller
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3×3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze differences in response distributions, clustering behavior, and alignment with self-identified stances. Our results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, while exhibiting differences in the resolution of belief structure and group separation. These findings suggest that LLM-generated surveys are suited for exploratory and large-scale analyses, and can be used to complement expert-designed instruments.
Where Is Politeness in Japanese BERT? A Layerwise Probing and CLS Activation Patching Study
Shusuke Hashimoto | Wenchen Shi
Shusuke Hashimoto | Wenchen Shi
Politeness is a central pragmatic dimension of language use, and Japanese honorifics offer a well-defined testbed for studying whether pretrained encoders represent socially meaningful distinctions. Prior BERT-based work has applied supervised models to Japanese honorific data, but we are not aware of analyses that localize honorific-level information across layers or test causal influence via activation patching in Japanese BERT-style encoders. We study these questions in LineDistilBERT using the KeiCO corpus, which labels sentences with four honorific levels. To isolate pretrained representations while still defining a task predictor, we freeze all encoder parameters and train only a lightweight [CLS] classification head as a minimal readout. We then run layerwise linear probing, training multinomial L2-regularized logistic-regression probes on [CLS] vectors from each layer to quantify linear decodability across depth and to select a best layer on development data. Finally, we test causal leverage with [CLS] activation patching, transplanting donor activations into receiver sentences at selected layers and measuring prediction transitions, logit shifts, and flip rates under standard controls. Overall, honorific level is broadly decodable across layers, and [CLS] interventions can systematically steer the frozen-encoder classifier with strong depth dependence, providing complementary evidence from probing and causal intervention for Japanese politeness in practice.
Rewrite the News: Tracing Editorial Reuse across News Agencies
Soveatin Kuntur | Nina Smirnova | Anna Wroblewska | Philipp Mayr | Sebastijan Razboršek Maček
Soveatin Kuntur | Nina Smirnova | Anna Wroblewska | Philipp Mayr | Sebastijan Razboršek Maček
This paper investigates sentence-level text reuse in multilingual journalism, analyzing where reused content occurs within articles. We present a weakly supervised method for detecting sentence-level cross-lingual reuse without requiring full translations, designed to support automated pre-selection to reduce information overload for journalists (Hołyst et al., 2024). The study compares English-language articles from the Slovenian Press Agency (STA) with reports from 15 foreign agencies (FA) in seven languages, using publication timestamps to retain the earliest likely foreign source for each reused sentence. We analyze 1,037 STA and 237,551 FA articles from two time windows (October 7–November 2, 2023; February 1–28, 2025) and identify 1,087 aligned sentence pairs after filtering to the earliest sources. Reuse occurs in 52% of STA articles and 1.6% of FA articles and is predominantly non-literal, involving paraphrase and compositional reuse from multiple sources. Reused content tends to appear in the middle and end of English articles, while leads are more often original, indicating that simple lexical matching overlooks substantial editorial reuse. Compared with prior work focused on monolingual overlap, we (i) detect reuse across languages without requiring full translation, (ii) use publication timing to identify likely sources, and (iii) analyze where reused material is situated within articles. Dataset and code: https://github.com/kunturs/lrec2026-rewrite-news.
OnCoCo 1.0: A Public Dataset for Fine-Grained Message Classification in Online Counseling Conversations
Jens Albrecht | Robert Lehmann | Aleksandra Poltermann | Eric Rudolph | Philipp Steigerwald | Mara Stieler
Jens Albrecht | Robert Lehmann | Aleksandra Poltermann | Eric Rudolph | Philipp Steigerwald | Mara Stieler
This paper presents OnCoCo 1.0, a new public dataset for fine-grained message classification in online counseling. It is based on a new, integrative system of categories, designed to improve the automated analysis of psychosocial online counseling conversations. Existing category systems, predominantly based on Motivational Interviewing (MI), are limited by their narrow focus and dependence on datasets derived mainly from face-to-face counseling. This limits the detailed examination of textual counseling conversations. In response, we developed a comprehensive new coding scheme that differentiates between 38 types of counselor and 28 types of client utterances, and created a labeled dataset consisting of about 2.800 messages from counseling conversations. We fine-tuned several models on our dataset to demonstrate its applicability. The data and models are publicly available to researchers and practitioners. Thus, our work contributes a new type of fine-grained conversational resource to the language resources community, extending existing datasets for social and mental-health dialogue analysis.
Predicting Social Media User Actions: A Hybrid Approach for Common and Rare Behavior Prediction on Bluesky
Benjamin White | Anastasia Shimorina
Benjamin White | Anastasia Shimorina
Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing approaches focus primarily on common actions like retweeting and liking, the prediction of rare but significant behaviors remains largely unexplored. This paper presents a hybrid methodology for social media user behavior prediction that addresses both frequent and infrequent actions across a diverse action vocabulary. We evaluate our approach on a large-scale Bluesky dataset containing 6.4 million conversation threads spanning 12 distinct user actions across 25 persona clusters. Our methodology combines four complementary approaches: (i) a lookup database system based on historical response patterns; (ii) persona-specific LightGBM models with engineered temporal and semantic features for common actions; (iii) a specialized hybrid neural architecture fusing textual and temporal representations for rare action classification; and (iv) generation of text replies. Our persona-specific models achieve an average macro F1-score of 0.64 for common action prediction, while our rare action classifier achieves 0.56 macro F1-score across 10 rare actions. These results demonstrate that effective social media behavior prediction requires tailored modeling strategies recognizing fundamental differences between action types. Our approach achieved first place in the SocialSim: Social-Media Based Personas challenge organized at the Social Simulation with LLMs workshop at the Conference on Language Modeling (COLM 2025).
The Data Acquisition Framework: Bridging Psychometrics and NLP for Personality Dataset Construction
Lorenz Dumanski | Michael Spranger | Melanie Siegel
Lorenz Dumanski | Michael Spranger | Melanie Siegel
Existing datasets for personality recognition in Natural Language Processing (NLP) suffer from documented quality problems: self-reported labels lacking psychometric validation, limited domain diversity and lack of context. Despite these known limitations, state-of-the-art approaches continue relying on the same datasets due to absence of alternatives. We present the Data Acquisition Framework (DAF), which addresses this gap by systematically translating psychometric questionnaire items into controlled communication scenarios through expert-community validation. DAF-items, validated scenario descriptions with contextual parameters, are deployed via the Automatic Data Acquisition and Annotation Tool (ADAAT). Participants complete personality surveys and engage in scenario-based text interactions with LLM personas configured to the DAF-Item context. This yields communication data with direct, item-level psychometric annotations.
Language Ideologies in a Multilingual Society: An LLM-based Analysis of Luxembourgish News Comments
Emilia Milano | Alistair Plum | Yves Scherrer | Christoph Purschke
Emilia Milano | Alistair Plum | Yves Scherrer | Christoph Purschke
Detecting language ideologies is a valuable yet complex task for understanding how identities are constructed through discourse. In Luxembourg’s multicultural and multilingual society, language ideologies reflect more than simple preferences: they carry deep cultural and social meanings, shaping identities and social belonging. Following recent developments in applying Natural Language Processing tools to linguistics and social science, this paper explores the potential of large language models to assist in the detection of language ideologies. We manually annotate a corpus of user comments in Luxembourgish with predefined ideological categories and then evaluate the performance of large language models under varying prompt conditions to assess their ability to replicate these human annotations. Since Luxembourgish is a small language and poorly represented in the LLMs’ training data, we also investigate whether machine-translating the data to high-resource languages increases performance on the ideology detection task. Our findings suggest that, while LLMs are not yet fully optimized for a multi-class ideological annotation task, they are practical tools to identify language ideological content.
Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents
Vahid Sadiri Javadi | Aksa Aksa | Fryderyk Karol Róg | Lucie Flek | Johanne Trippas
Vahid Sadiri Javadi | Aksa Aksa | Fryderyk Karol Róg | Lucie Flek | Johanne Trippas
Social interactions are shaped by the interplay of dispositional traits and situational context, yet systematically investigating how personality configurations between individuals jointly influence social behavior across diverse social contexts remains methodologically challenging. We address this gap by introducing a simulation pipeline adapted from the CHARISMA framework, which employs well-known movie characters and public figures as psychologically grounded agents for multi-LLM social simulation using a method we term personality anchoring. We present a large-scale empirical study examining how dyadic Agreeableness composition influences social interaction outcomes across 1,010 simulated conversations. Our results reveal a monotonic relationship between dyadic Agreeableness composition and shared goal achievement, with Homogeneous-Agreeable pairs achieving success 10 times the rate of Homogeneous-Disagreeable pairs (62% vs. 6%). Behavioral mediation analysis reveals that Agreeableness shapes goal achievement partially through cooperative strategy selection, though it continues to predict outcomes within the same dominant strategy, indicating pathways beyond observable conversational behavior. Robustness analyses confirm high consistency of results across repeated simulations (ICC = 0.89) and stable personality expression across diverse scenarios, validating personality anchoring as a viable operationalization strategy.