Johan Sjons
2026
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
Nedjma Ousidhoum | Junho Myung | Carla Perez-Almendros | Jiho Jin | Amr Keleg | Meriem Beloucif | Yi Zhou | Rodrigo Agerri | Vladimir Araujo | Naomi Baes | James Barry | Joanne Boisson | Nancy F. Chen | Christine de Kock | Aleksandra Edwards | Joseba Fernandez de Landa | Mohamed Fazli Imam | Huda Hakami | Shu-Kai Hsieh | Joseph Marvin Imperial | Roy Ka-Wei Lee | Zhengyuan Liu | Chenyang Lyu | Younes Samih | Johan Sjons | Bryan Tan | Asahi Ushio | Weihua Zheng | Alice Oh | Jose Camacho-Collados
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Nedjma Ousidhoum | Junho Myung | Carla Perez-Almendros | Jiho Jin | Amr Keleg | Meriem Beloucif | Yi Zhou | Rodrigo Agerri | Vladimir Araujo | Naomi Baes | James Barry | Joanne Boisson | Nancy F. Chen | Christine de Kock | Aleksandra Edwards | Joseba Fernandez de Landa | Mohamed Fazli Imam | Huda Hakami | Shu-Kai Hsieh | Joseph Marvin Imperial | Roy Ka-Wei Lee | Zhengyuan Liu | Chenyang Lyu | Younes Samih | Johan Sjons | Bryan Tan | Asahi Ushio | Weihua Zheng | Alice Oh | Jose Camacho-Collados
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
We present our shared task on evaluating the adaptability of LLMs and NLP systems across multiple languages and cultures. The task data consist of an extended version of our manually constructed BLEnD benchmark (Myung et al., 2024), covering more than 30 language–culture pairs, predominantly representing low-resource languages spoken across multiple continents. As the task is designed strictly for evaluation, participants were not permitted to use the data for training, fine-tuning, few-shot learning, or any other form of model modification.Our task includes two tracks: (a) Short-Answer Questions (SAQ) and (b) Multiple-Choice Questions (MCQ). Participants were required to predict labels and were allowed to submit any NLP system and adopt diverse modelling strategies, provided that the benchmark was used solely for evaluation. The task attracted more than 140 registered participants, and we received final submissions from 62 teams, along with 19 system description papers.We report the results and present an analysis of the best-performing systems and the most commonly adopted approaches. Furthermore, we discuss shared insights into open questions and challenges related to evaluation, misalignment, and methodological perspectives on model behaviour in low-resource languages and for under-represented cultures. Our data and resources are available at https://github.com/BLEnD-SemEval2026/SemEval-2026-Task-7.
Cultural Grounding in Swedish: Extending an Everyday Knowledge Benchmark for LLMs
Meriem Beloucif | Johan Sjons
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Meriem Beloucif | Johan Sjons
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Benchmarks for evaluating Large Language Models (LLMs) on everyday knowledge across cultures and languages are increasingly used to assess cultural competence and contextual understanding. However, many multilingual extensions rely primarily on translated question–answer pairs, limiting their ability to capture locally grounded variation. In this work, we present a Swedish extension of an existing cross-cultural everyday knowledge benchmark in which questions are translated into Swedish, and answers are individually collected from five participants, coming from diverse social and professional backgrounds. This design enables us to capture culturally situated, naturally produced responses rather than transferred or translated answer templates. We document the translation protocol, annotators, and agreement analysis, and examine variation across annotators as a signal of culturally contingent knowledge. We evaluate several state-of-the-art multilingual and instruction-tuned LLMs against the aggregated human responses and analyze model performance. Our results reveal that while models often approximate prototypical answers, they struggle with culturally specific nuances and intra-cultural variation. The Swedish extension provides a resource for studying culturally grounded evaluation and highlights the importance of human-generated local answers when benchmarking LLMs across languages.
The Swedish Benchmark of Linguistic Minimal Pairs
Johan Sjons | Fredrik Heinat | Murathan Kurfali
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Johan Sjons | Fredrik Heinat | Murathan Kurfali
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We introduce the Swedish Benchmark of Linguistic Minimal Pairs, a dataset for evaluating syntactic performance in language models. It includes 2,500 minimal pairs organized into 25 syntactic phenomena, with 100 pairs per phenomenon. Each pair contrasts a well-formed and an ill-formed sentence that differ minimally. For each phenomenon, we manually constructed ten pairs from scratch. We semi-automatically generated the remaining 90 pairs and manually adjusted them. A random sample was assessed by 40 participants, who selected the well-formed sentence in 98.05% of cases. We evaluate eleven state-of-the-art models. Results generally show that models handle local agreement well but struggle with certain long-distance dependencies and word order phenomena. Model size seems to matter less than the training domain. Prompt-based evaluation generally lowers performance. We show that model performance is stable across handcrafted and generated subsets and across sample sizes, suggesting that 100 pairs per phenomenon suffice for reliable evaluation. Future work will expand the number of phenomena.
How Much Data for Stable Formant Values? Pipeline for Convergence Detection Based on Read Speech
Kayla Sward | Johan Sjons | Axel G. Ekstrom
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Kayla Sward | Johan Sjons | Axel G. Ekstrom
Proceedings of the Fifteenth Language Resources and Evaluation Conference
This study investigates the stability and convergence of vowel formants (F1, F2, F3) in read speech through an extensive corpus of audiobook recordings. While most formant studies rely on brief, isolated utterances recorded in laboratory settings, this analysis draws on 3,384 chapters (about 942 hours) of continuous, stylistically varied speech from publicly available audiobooks. The data was processed using an automated pipeline that comprised transcription, phoneme alignment, and formant extraction. Several statistical techniques – First Token Within (FTW), Cumulative Sum (CUSUM), Two-Sample t-Test, Confidence Interval (CI) Shrinkage, Piecewise Linear Fitting (PWLF), and Binary Segmentation (BinSeg) – were compared for their effectiveness in identifying stabilization points. Findings indicate that formant means generally stabilize within 60 to 230 vowel tokens per phoneme, dependent on vowel type and speaker gender. Of the methods that were evaluated, CUSUM yielded the most consistent and informative results. The results provide practical guidelines for determining the quantity of non-laboratory speech required to obtain reliable vowel formant averages.
Neural Network-assisted Analysis of Tube Vocal Tract Models
Runhui Song | Johan Sjons | Axel G. Ekstrom
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Runhui Song | Johan Sjons | Axel G. Ekstrom
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We present a pipeline for deep neural network assisted modeling and analysis of the behavior of an acoustic tube. The vocal tract is represented as a series of cylindrical tube segments, each characterized by fixed length and variable cross-sectional area. A large synthetic dataset of such tube configurations is generated, and a circuit theory–based algorithm predicts corresponding formant frequencies. To explore mapping between vocal tract shapes and formant values, the pipeline integrates both linear regression and nonlinear machine learning models - including multilayer perceptrons. Model interpretability is measured using Shapley Additive Explanations (SHAP), which quantifies the contribution of each segment to predicted formant frequencies. The proposed framework enables detailed exploration of the articulatory-acoustic relationships inherent to an acoustic tube and vocal tract simulacrum. We present and describe the pipeline in the context of modeling effects of perturbations on the first three formants for a 16-cm tube, divided into 1 cm segments. Our pipeline can be applied to any method that models predictions of behavior of an acoustic tube, where the tube is conceived as a series of segmented units.
Do large language models and humans follow similar learning stages? Assessing GPT-2’s order of Swedish grammar acquisition within the Processability Theory framework
Stella Lundqvist | Murathan Kurfali | Johan Sjons
Proceedings of the 1st Workshop on Computational Developmental Linguistics (CDL)
Stella Lundqvist | Murathan Kurfali | Johan Sjons
Proceedings of the 1st Workshop on Computational Developmental Linguistics (CDL)
We investigate whether GPT-2 acquires Swedish grammatical structures in the same implicational order as for human second language (L2) learners, as predicted by Processability Theory (PT). We present SwePT – a minimal pair dataset targeting Swedish syntactic and morphological structures that are acquired by human L2 learners on four separate stages of language development – and evaluate the GPT-2 models on SwePT using an acceptability classification task throughout fine-tuning with different input orders in regards to the grammatical structures identified in the data. We find that the observed acquisition orders correlate across the fine-tuned models, while violating the implicational order sequence as hypothesized by PT. The observed relation between performance on the classification task and frequency distributions of the contrasting features in the minimal pairs suggests that the acquisition order can be explained by unigram and n-gram heuristics. While the adaptation of NLP methodologies into the PT framework requires further conceptual and methodological refinement, we do not find evidence for PT-like grammatical development in our experiments.
2020
A Multi-word Expression Dataset for Swedish
Murathan Kurfalı | Robert Östling | Johan Sjons | Mats Wirén
Proceedings of the Twelfth Language Resources and Evaluation Conference
Murathan Kurfalı | Robert Östling | Johan Sjons | Mats Wirén
Proceedings of the Twelfth Language Resources and Evaluation Conference
We present a new set of 96 Swedish multi-word expressions annotated with degree of (non-)compositionality. In contrast to most previous compositionality datasets we also consider syntactically complex constructions and publish a formal specification of each expression. This allows evaluation of computational models beyond word bigrams, which have so far been the norm. Finally, we use the annotations to evaluate a system for automatic compositionality estimation based on distributional semantics. Our analysis of the disagreements between human annotators and the distributional model reveal interesting questions related to the perception of compositionality, and should be informative to future work in the area.
Search
Fix author
Co-authors
- Murathan Kurfali 3
- Meriem Beloucif 2
- Axel G. Ekstrom 2
- Rodrigo Agerri 1
- Vladimir Araujo 1
- Naomi Baes 1
- James Barry 1
- Joanne Boisson 1
- Jose Camacho-Collados 1
- Nancy Chen 1
- Aleksandra Edwards 1
- Mohamed Fazli Imam 1
- Joseba Fernandez de Landa 1
- Huda Hakami 1
- Fredrik Heinat 1
- Shu-Kai Hsieh 1
- Joseph Marvin Imperial 1
- Jiho Jin 1
- Amr Keleg 1
- Roy Ka-Wei Lee 1
- Zhengyuan Liu 1
- Stella Lundqvist 1
- Chenyang Lyu 1
- Junho Myung 1
- Alice Oh 1
- Nedjma Ousidhoum 1
- Carla Perez-Almendros 1
- Younes Samih 1
- Runhui Song 1
- Kayla Sward 1
- Bryan Tan 1
- Asahi Ushio 1
- Mats Wirén 1
- Weihua Zheng 1
- Yi Zhou 1
- Christine de Kock 1
- Robert Östling 1