Shu-Kai Hsieh
Papers on this page may belong to the following people: Shu-Kai Hsieh, Shu-Kai HSIEH
2026
Lexical Familiarity Predicts Processing Depth for Nonliteral Language in Large Language Models
Lang-Ching Yeh | Yu-Chieh Wang | Shu-Kai Hsieh
Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026)
Lang-Ching Yeh | Yu-Chieh Wang | Shu-Kai Hsieh
Proceedings of the 6th Workshop on Trustworthy NLP (TrustNLP 2026)
This paper investigates how large language models internally process nonliteral language. Analyzing five categories spanning slang, metaphor, and idioms across all 48 layers of Gemma-3-12B-IT with Gemma Scope 2 sparse autoencoders, we find a lexical familiarity gradient: processing depth depends on available prior lexical knowledge, not figurative type. Idioms diverge at L1 as entrenched units; expressions built from familiar words (metaphors, semantic-shift and constructional slang) converge at L7–9; neologisms peak at L41, activating 3× more unique features. Paraphrase residual analysis confirms strong signals only at the gradient endpoints, yielding a three-tier hierarchy of entrenched retrieval, known-word reanalysis, and novel-word construction. Crucially, this peak-layer structure replicates in base models (Gemma-PT, Qwen-Base), demonstrating that the gradient is a robust property of pretrained representations rather than an alignment artifact. We additionally identify an activation density confound in SAE feature counts that produces spurious cross-condition convergence. Overall, processing depth is better predicted by lexical familiarity than by figurative type, with implications for robustness to non-standard language and for SAE-based interpretability.
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Nedjma Ousidhoum | Junho Myung | Carla Perez-Almendros | Jiho Jin | Amr Keleg | Meriem Beloucif | Yi Zhou | Rodrigo Agerri | Vladimir Araujo | Naomi Baes | James Barry | Joanne Boisson | Nancy F. Chen | Christine de Kock | Aleksandra Edwards | Joseba Fernandez de Landa | Mohamed Fazli Imam | Huda Hakami | Shu-Kai Hsieh | Joseph Marvin Imperial | Roy Ka-Wei Lee | Zhengyuan Liu | Chenyang Lyu | Younes Samih | Johan Sjons | Bryan Tan | Asahi Ushio | Weihua Zheng | Alice Oh | Jose Camacho-Collados
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Nedjma Ousidhoum | Junho Myung | Carla Perez-Almendros | Jiho Jin | Amr Keleg | Meriem Beloucif | Yi Zhou | Rodrigo Agerri | Vladimir Araujo | Naomi Baes | James Barry | Joanne Boisson | Nancy F. Chen | Christine de Kock | Aleksandra Edwards | Joseba Fernandez de Landa | Mohamed Fazli Imam | Huda Hakami | Shu-Kai Hsieh | Joseph Marvin Imperial | Roy Ka-Wei Lee | Zhengyuan Liu | Chenyang Lyu | Younes Samih | Johan Sjons | Bryan Tan | Asahi Ushio | Weihua Zheng | Alice Oh | Jose Camacho-Collados
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al., 2024), covering more than 30 language–culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification.Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers.We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures. Our data and resources are available at https://github.com/BLEnD-SemEval2026/SemEval-2026-Task-7.
Search
Fix author
Co-authors
- Rodrigo Agerri 1
- Vladimir Araujo 1
- Naomi Baes 1
- James Barry 1
- Meriem Beloucif 1
- Joanne Boisson 1
- Jose Camacho-Collados 1
- Nancy Chen 1
- Aleksandra Edwards 1
- Mohamed Fazli Imam 1
- Joseba Fernandez de Landa 1
- Huda Hakami 1
- Joseph Marvin Imperial 1
- Jiho Jin 1
- Amr Keleg 1
- Roy Ka-Wei Lee 1
- Zhengyuan Liu 1
- Chenyang Lyu 1
- Junho Myung 1
- Alice Oh 1
- Nedjma Ousidhoum 1
- Carla Perez-Almendros 1
- Younes Samih 1
- Johan Sjons 1
- Bryan Tan 1
- Asahi Ushio 1
- Yu-Chieh Wang 1
- Lang-Ching Yeh 1
- Weihua Zheng 1
- Yi Zhou 1
- Christine de Kock 1