Proceedings of the Workshop Neology and Large Language Models

Giedre Valunaite Oleskeviciene, Voula Giouli, Florentina Armaselu, Chaya Liebeskind, Barbara McGillivray (Editors)



We present a scalable, modular pipeline for automatic neologism detection that combines rule-based filtering with LLM classification. The pipeline is grounded in two complementary word-formation frameworks, grammatical and extra-grammatical morphology, which jointly define the scope of what counts as a neologism and inform a four-class classification scheme (NEOLOGISM, ENTITY, FOREIGN, NONE). While designed to be modular and transferable at the architectural level, the pipeline is instantiated on 527 million English-language Reddit posts spanning 2005–2024. From this corpus, we extract 124.6 million unique tokens and reduce them by over 99.99% to yield 1,021 neologism candidates, a set small enough for manual expert verification. Multiple LLMs independently classify each candidate via majority vote, with a final verification step, revealing substantial cross-model disagreement and highlighting the challenge of operationalizing neologism detection at scale. Manual annotation of all 1,021 candidates confirms that 599 (58.7%) are genuine lexical innovations.
Large language models (LLMs) are increasingly deployed to detect, generate, and normalize neologisms across languages. While prior work has examined their capacity to model semantic change and handle temporal drift, insufficient attention has been paid to how training-data asymmetries interact with probabilistic generation mechanisms to structure lexical innovation itself. This paper argues that AI-driven neology is shaped by systematic high-resource bias that privileges dominant languages in the production, stabilization, and dissemination of new lexical items. Drawing on sociolinguistics, language political economy, lexicography, and computational modeling theory, we formalize how distributional imbalance alters innovation likelihood across languages. We introduce a taxonomy of bias types specific to AI-mediated neology, present a probabilistic account of generative reinforcement loops, and illustrate these mechanisms using documented examples from English-Arabic and English-Icelandic language pairs. We derive empirically testable predictions and propose concrete mitigation strategies for lexicographers, language planners, and NLP researchers.
Large language models (LLMs) are increasingly used for writing assistance in small contact languages, yet it is unclear whether they respect community norms around lexical borrowing and neology. We introduce LexNeo-Bench, a 3,050-instance token-level benchmark derived from LuxBorrow, a large-scale Luxembourgish news corpus, where target tokens are labelled as native or as French, German, or English borrowings. Using this benchmark, we probe three multilingual LLMs across 34 prompt settings on two tasks: borrowing type classification and a binary lexical-innovation proxy (borrowing versus native). Without external context, models perform only slightly above chance on borrowing classification, so we construct a linguistic knowledge graph that encodes donor language, morphological patterns, and lexical analogues, and inject instance-specific subgraphs into the prompt. Knowledge-graph prompts raise borrowing classification accuracy from 25 – 35% up to 71 – 81% and largely close the gap between small and large models, while leaving neology detection difficult and sensitive to few-shot design. Our results show that lexicon-aware prompting is highly beneficial for robust borrowing judgments in low-resource contact languages and that lexical resources can serve as structured context for LLM evaluation. This study was carried out within the ENEOLI COST Action and examines borrowing as a form of lexical innovation in multilingual Luxembourgish data.
Lexical innovation refers to the process of creating new lexical items, enabling languages to adapt to evolving socio-cultural and material realities. The domains of business, economics, and finance are among the most productive ones of lexical innovation. The present research study lies at the intersection of lexical innovation, idiomaticity, and large language model (henceforth, LLM) research and investigates lexical productivity, semantic shift, and globalization (Anglocentric changes) in business-related colour idioms by comparing human translation and annotation with the output of LLMs. The current experiment involves an initial study carried out for five languages: English (the pivotal one), Albanian (AL), Hebrew (HE), Hungarian (HU), Lithuanian (LT), and Standard European Portuguese (PT). The research results reveal that LLMs show high mutual agreement, but the agreement with humans is lower. The internal consistency of LLMs reflects shared Anglocentric metaphor encoding rather than convergence toward human idiomatic usage. It demonstrates that human expertise remains essential for high-quality idiomatic translation, particularly for culture-specific expressions.
English neologisms, or newly coined words, have previously been shown to emerge in sparser semantic neighborhoods (filling semantic gaps) and near other neologisms (in growing semantic areas). In this work, we investigate where in semantic space Spanish neologisms emerge, and whether this mirrors English neologism development. We find that Spanish neologisms, in comparison to non-neologisms, do indeed appear both nearer to other neologisms and further from non-neologisms. We additionally investigate the prevalence of loanwords from other languages through time in Spanish neologism production and manually assess the topics that appear as loanwords at four years: 1810, 1900, 1950, and 1990. Our findings show that on average, the Spanish neologisms in our dataset have fewer neighboring words in semantic space compared to non-neologisms and tend to cluster more tightly in the semantic space, indicating that patterns of neologism emergence span languages. This suggests that novel methods for neologism detection may be cross-lingually applicable, with these features serving as multilingual predictors of neologism emergence.
The English language is changing faster than before, partly due to the influence of the Internet. Digital language includes a large number of discourse markers (DMs), many of which can be considered innovative. Acronymization, pragmatic specialisation, and compensatory lexical innovation are the most common lexical processes that can be witnessed in the DMs used in computer-mediated communication (CMC). The following novel DMs were identified in recent Twitter chats: lol, tbh, omg, meh, and idk. These DMs perform several functions, such as showing emotions, signalling uncertainty, hesitation, or mitigation. Interpreting these functions may not be an easy or obvious task for AI. The primary aim of the study is to evaluate the pragmatic competence of an LLM, Gemini 3 Pro, regarding the interpretation of these novel DMs. A mixed-method research process was employed: LLM-generated outputs were compared with the findings of the relevant literature, quantitative corpus analysis, and our qualitative human interpretation to assess the model’s analytical usefulness. Gemini 3 Pro was found to show a high level of pragmatic competence in terms of interpreting the functions of DMs, but sometimes tended to overgeneralise, or failed to understand the tone of the text and the intention of the speaker to use a DM.
The growing popularity and misconceptions about conversational AI systems are driving efforts to establish a universally accepted framework for evaluating large language models. Testing large language models on tasks designed to assess human cognitive skills has become widespread. This paper presents the results of a pilot experiment and a comparative evaluation of the ability of OpenAI’s GPT-4.1 and GPT-4.1 mini to detect semantic ambiguity based on the works of Shultz and Pilon (1973) and Zipke et al. (2009). The experiment used a task sheet of 116 items utilising riddles, single sentences, and sentence pairs. It included systematically varied instructions on a four-level scale ranging from no mention of ambiguity to direct mention. Lexical and structural ambiguity were both employed, including surface-structure and deep-structure ambiguity. The results suggest that even advanced models, such as GPT-4.1 and GPT-4.1 mini, tend to consider only one possible meaning of ambiguous sentences. However, the recognition of ambiguity improved quickly when the possibility of ambiguity was explicitly referenced in the instruction. Additionally, the results imply that model size is not directly connected to performance, as GPT-4.1 scored better on lexical ambiguity detection tasks, while GPT-4.1 mini surpassed the larger model in structural ambiguity detection. The findings prove that future research with a more complex experiment design based on the same principles would be beneficial.
Large language models (LLMs) are increasingly used for lexicographic support, neology detection, and semantic categorization, yet their behaviour on historical newspapers remains under-evaluated. This short paper describes an ongoing project that extends a DH2026-accepted two-phase methodology for extracting and tracking rumours in historical newspapers. From large-scale US and UK corpora (PleIAs/US-PD-Newspapers; biglam/hmd_newspapers), the DH workflow produces gold-standard sentence-level rumour instances with proposition-like “rumour content” spans (Rumour_Content/Cleaned_Content) and extraction-pattern metadata. Building on these historically grounded units, we propose an LLM-centered benchmark and analysis pipeline for assigning topical frames and evidential stance to rumour propositions, and for auditing “temporal projection” when models introduce anachronistic modern misinformation framings. For controlled cross-variety comparison we construct a strictly balanced benchmark of 800 instances over two well-attested bins (1840–1859, 1860–1879) and both national varieties (200 per country per bin). We outline prompt conditions (text-only vs time-aware vs historically calibrated) and self-consistency voting to quantify label stability and error modes. A small manually annotated subset supports evaluation, while the main contribution is the benchmark design, prompts, and reproducible protocol enabling community feedback before full-scale results are finalized.