Urban Knupleš
Author directory2026
Literally Concrete or Figuratively Abstract? Multilingual Concreteness Norms for Verb-Object Expressions
Urban Knupleš | Diego Frassinelli | Alexander Fraser | Sabine Schulte im Walde
Transactions of the Association for Computational Linguistics, Volume 14
Urban Knupleš | Diego Frassinelli | Alexander Fraser | Sabine Schulte im Walde
Transactions of the Association for Computational Linguistics, Volume 14
While existing concreteness norms primarily target words in isolation, little attention has been paid to concreteness in context. To address this, we systematically collect multilingual concreteness ratings using Best-Worst Scaling (BWS) for 5,814 verb-direct object noun expressions in three languages with different degrees of resource availability: English, German, and Slovene. We identify consistent patterns where the concreteness of verb-noun combinations is more strongly influenced by the nominal object than the verb. Through comparative analyses on an English subset, we demonstrate that BWS guarantees more reliable concreteness judgments than traditional rating scales. Expanding beyond our human-generated data, we use traditional and LLM-based automatic extrapolation methods to generate a large-scale multilingual resource of over 430,000 expressions. Additionally, we conduct a study examining the interaction between concreteness and literal vs. figurative judgments for a subset of 1,800 expressions in all three languages, along with example usage sentences. Our findings show that lower concreteness ratings correlate with figurative language, thus reinforcing the link between abstractness and figurativeness. All resources are available from https://github.com/urbikn/multilingual-concreteness-vo.
2024
Gender Identity in Pretrained Language Models: An Inclusive Approach to Data Creation and Probing
Urban Knupleš | Agnieszka Falenska | Filip Miletić
Findings of the Association for Computational Linguistics: EMNLP 2024
Urban Knupleš | Agnieszka Falenska | Filip Miletić
Findings of the Association for Computational Linguistics: EMNLP 2024
Pretrained language models (PLMs) have been shown to encode binary gender information of text authors, raising the risk of skewed representations and downstream harms. This effect is yet to be examined for transgender and non-binary identities, whose frequent marginalization may exacerbate harmful system behaviors. Addressing this gap, we first create TRANsCRIPT, a corpus of YouTube transcripts from transgender, cisgender, and non-binary speakers. Using this dataset, we probe various PLMs to assess if they encode the gender identity information, examining both frozen and fine-tuned representations as well as representations for inputs with author-specific words removed. Our findings reveal that PLM representations encode information for all gender identities but to different extents. The divergence is most pronounced for cis women and non-binary individuals, underscoring the critical need for gender-inclusive approaches to NLP systems.
2023
Investigating the Nature of Disagreements on Mid-Scale Ratings: A Case Study on the Abstractness-Concreteness Continuum
Urban Knupleš | Diego Frassinelli | Sabine Schulte im Walde
Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)
Urban Knupleš | Diego Frassinelli | Sabine Schulte im Walde
Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL)
Humans tend to strongly agree on ratings on a scale for extreme cases (e.g., a CAT is judged as very concrete), but judgements on mid-scale words exhibit more disagreement. Yet, collected rating norms are heavily exploited across disciplines. Our study focuses on concreteness ratings and (i) implements correlations and supervised classification to identify salient multi-modal characteristics of mid-scale words, and (ii) applies a hard clustering to identify patterns of systematic disagreement across raters. Our results suggest to either fine-tune or filter mid-scale target words before utilising them.