2026

While large language models (LLMs) are widely used across cultures, they often generate culturally inappropriate responses in unfamiliar cultural contexts due to biases embedded in their training data. Existing approaches primarily rely on expanding static cultural knowledge, which fails to capture the inherently relative and context-dependent nature of culture. In this paper, we propose a Cultural value-based Reasoning (CURE) framework that interprets behaviors through underlying cultural value systems. In addition, we integrate CURE into LLMs via Chain-of-Thought (CoT) distillation, referred to as CURE-distillation, to internalize culturally grounded reasoning. Experimental results show that models trained with CURE-distillation improve cultural adaptability, enabling them to produce culturally aligned ethical judgments across diverse cultural scenarios. These results suggest that strengthening sociocultural reasoning capabilities can substantially improve the cultural adaptability of LLMs. The code is available at https://github.com/KUNLP/CURE.

2022

CODI-CRAC 2022 Shared Task in Dialogues consists of three sub-tasks: Sub-task 1 is the resolution of anaphoric identity, sub-task 2 is the resolution of bridging references, and sub-task 3 is the resolution of discourse deixis/abstract anaphora. Anaphora resolution is the task of detecting mentions from input documents and clustering the mentions of the same entity. The end-to-end model proceeds with the pruning of the candidate mention, and the pruning has the possibility of removing the correct mention. Also, the end-to-end anaphora resolution model has high model complexity, which takes a long time to train. Therefore, we proceed with the anaphora resolution as a two-stage pipeline model. In the first mention detection step, the score of the candidate word span is calculated, and the mention is predicted without pruning. In the second anaphora resolution step, the pair of mentions of the anaphora resolution relationship is predicted using the mentions predicted in the mention detection step. We propose a two-stage anaphora resolution pipeline model that reduces model complexity and training time, and maintains similar performance to end-to-end models. As a result of the experiment, the anaphora resolution showed a performance of 68.27% in Light, 48.87% in AMI, 69.06% in Persuasion, and 60.99% on Switchboard. Our final system ranked 3rd on the leaderboard of sub-task 1.