Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)

John P. McCrae, Katerina Gkirtzou, Fahad Khan, Patricia Martin Chozas, Sara Carvalho, Erin Canning (Editors)



Parliamentary discourse constitutes a key domain in which political actors publicly articulate policy positions and priorities through language. This study investigates debates from the Italian Chamber of Deputies (1948–2006) to identify and analyse latent semantic themes and their evolution using BERTopic-based dynamic topic modeling. The analysis relies on a subset of the ItaParlCorpus (Cova, 2025), a large-scale, machine-readable corpus enriched with temporal, institutional, and political metadata. Beyond topic extraction,this work addresses a largely unexplored challenge: the formalization of topics derived from unsupervised, embedding-based topic modeling as Linked Data entities, adopting a linguistic perspective. Extracted topics are formalized as semantic entities reusing the OntoLex–Lemon model, its FrAC extension and declaring a dedicated ontology to link topics to speeches, speakers, political parties, and temporal information reusing standardized vocabularies and persistent URIs. This integration enables semantic querying through SPARQL, supporting analyses of topic distributions across political actors, parties and illustrating the analytical potential of the proposed approach. Moreover, the study highlights limitations in the formalization of topic modeling outputs, particularly regarding the representation of ambiguous word forms and their alignment with lexical concepts in OntoLex–Lemon.
This paper presents the first core component of LinkEn, a knowledge base of interoperable language resources for English adhering to Linked Open Data principles. With this initial step towards a broader infrastructure, we focus on the development of a lemma-centered hub designed to enable interoperability between distributed lexical resources, corpora, and linguistic annotations. The modeling is inspired by the LiLa Knowledge Base for Latin and the OntoLex-Lemon model, ensuring compatibility with existing lemma-centric knowledge graphs and enabling future cross-linguistic interoperability. Rather than relying solely on manual knowledge graph construction and significant human effort, the lemma bank has been developed through a hybrid neuro-symbolic pipeline that integrates large language models into the generation of RDF data under explicit ontological constraints. This approach combines automated generation with ontology-driven supervision and evaluation, enabling scalable yet controlled construction of structured lexical knowledge. By presenting the first steps towards the LinkEn Knowledge Base, this paper contributes both a new lemma bank for English and an experimental methodology for the semi-automatic creation of Linked Data based knowledge graphs.
We present an ongoing effort to bridge the Lessico dei Beni Culturali (LBC), a multilingual lexicographic project cov- ering Italian cultural heritage terminology, with the Linguistic Linked Open Data (LLOD) ecosystem. The LBC corpus spans five centuries of art-historical writing, from fifteenth- and sixteenth-century treatises by Alberti, Leonardo, and Vasari to nineteenth-century works by Stendhal and Burckhardt and contemporary tourist guides to Florence, with source texts in several European languages alongside their translations. The resource has already undergone automatic linguistic annotation and term extraction, but lacks structured lexical representation in any standard LLOD formalism. We describe the current state of the resource, identify the main challenges for its publication as Linked Data — including the modelling of culturally-bound terms (realia), historical proper nouns, and multilingual source texts of different registers — and outline a roadmap towards its representation in OntoLex-Lemon (McCrae et al., 2017) and its alignment with existing LLOD resources such as the Getty Vocabularies (Getty Research Institute, 2024a) and Wikidata (Vrandecic and Krötzsch, 2014). By sharing this work with the LLOD community, we expect input on best practices for historical-artistic and cultural heritage lexicons that will raise interoperability between resources from different sources, generating new information and increasing the value of existing data.
The humanities are a vast and highly diverse field – both methodologically and technologically –, so, it is not unsurprising to see independent researchers or projects to work on the same data, and producing complementary, but technically incompatible electronic editions from the same source material. We suggest that existing Linguistic Linked Open Data (LLOD) technology can play a crucial role for performing a post-hoc consolidation of their efforts, illustrated for the Old Saxon (Old Low German) Heliand, a 9th c. gospel harmony previously annotated for different aspects of syntax in three independent research projects and over different versions (editions and manuscripts) of the original text. We describe the derivation of a UD-compliant corpus from the consolidation of the existing annotations. This includes the transformation of the original annotations to corpus-specific CoNLL (TSV) formats, the alignment between the different corpora, and their integration. A particular challenge is the processing of incomplete annotations, as one of the source corpora (Heliand B4) provides non-recursive nominal and clausal chunks only, and another corpus (Heliand DDD) even only sentence boundaries, clause types and parts of speech, but no actual phrasal structures. In this paper, we specifically focus on the application of Fintan (CoNLL-RDF) and SPARQL for performing the necessary graph rewriting operations.
Our understanding of social reality is shaped by the specific ways in which that reality is framed by different sources. Analyzing framing means examining how these sources are able to convey particular worldviews by foregrounding or downplaying certain aspects of experience. Current computational approaches address this task by automatically identifying communicative patterns (e.g., topic selection or rhetorical strategies) that characterize individual artifacts. However, they often remain document-bound, overlooking the comparative dimension that enables the uncovering of convergent or conflicting narratives about the same actor, event, or issue. In this paper, we propose DORIS, an ontology that supports both document-level and cross-document framing analysis using SPARQL queries on automatically constructed Knowledge Graphs. We validate the proposed approach through a case study of historical news articles, exploring multiple framings of a real-world event using Fillmore’s Frame Semantics and the FrameNet resource. Code and data are available on GitHub at https://github.com/beatrice-f/DORIS/.
This paper presents LaReS (Latin Represented Speech), a Linked Open Data resource designed to model represented speech in Latin literature and to align the DICES database of direct speeches in Greek and Latin epic with the LiLa Knowledge Base. While DICES provides a rich collection of metadata on direct speech in epic poetry, its operational approach and its relatively shallow conceptual modeling limit its interoperability and extensibility. The modeling strategy implemented in LaReS is based on the separation of the textual level from the narratological dimension. CIDOC CRM and DOLCE+DnS are used to conceptualize the basic notions in the two modules. LaReS now includes 341 speeches in Virgil’s Aeneid, linking 36,782 tokens in LiLa to speech units derived from DICES
We present Open English NameNet, a new large-scale lexical resource that extends Open English Wordnet with named entities derived from Wikidata. While English Wordnet has historically included many proper nouns, its coverage has been incomplete and inconsistent, and encyclopedic knowledge sources have grown rapidly in parallel. To address this gap, we systematically extract and align named entities from Wikidata with the Open English Wordnet hierarchy, ensuring each entity is appropriately placed through instance hypernym relations. Our methodology combines existing WordNet–Wikipedia mappings with Wikidata information and applies domain-specific strategies for people, plants and animals, and languages, to account for structural and semantic differences between the resources. This approach results in the largest English lexical-semantic resource currently available, with extensive coverage and structured integration. We release the resource openly to support the development of lexically and encyclopedically informed language technologies.
This paper presents the requirements and implementation details for a new core module of the OntoLex-Lemon model, representing the first major evolution of the de-facto standard since its 2016 release. While the original model successfully bridged ontologies and dictionaries through “semantics by reference,” community adoption has identified critical gaps in handling lexicographic structures and retrodigitized resources. We detail a community-driven methodology that identified fifteen key requirements and we present a proposed architecture for the new OntoLex core, which integrates elements from the Lexicography module and addresses both semantic web and lexicography use cases. Further, we improve interoperability with standards like DMLex and TEI-Lex0 while maintaining strict backwards compatibility for existing users of the model.
In this paper, we present the NILOMORPH project, that aims at describing the complex non-concatenative morphology of West Nilotic languages and reconstructing the dynamics of its evolution from a more straightforward concatenative system. The project adopts techniques from several methodologies and draws on many kinds of data displaying different formats, tagsets and conventions. Data are also multilingual, documenting different West Nilotic varieties, and multimodal, including also audio and video recordings. This makes the process of integration of these data particularly challenging. We first describe how the data can be converted to standard formats such as CLDF and Paralex, to achieve interoperability between resources of the same kind. We then discuss how they can be modelled as Linguistic Linked Open Data in the Resource Description Framework, reusing already existing vocabularies and defining new classes and properties to meet the needs of the project, to also achieve interoperability between resources of different kinds.
This paper introduces the Research Constructicon (RCxn), a project developed within the Research Training Group Dimensions of Constructional Space. The training group finances PhD projects in the framework of Construction Grammar (CxG), which views language as a network of form-meaning pairings. The RCxn is designed as a dynamic, community-driven resource that documents linguistic constructions while also capturing the research processes and findings associated with them. The project addresses three core dimensions: (1) the development of a modular ontology to represent constructions, their relationships, and the research surrounding them; (2) the implementation of database populated by researchers’ contributions; and (3) the creation of a web application to visualize and interact with the data. This paper focuses on our work to implement a rich ontology for the RCxn, which has to accommodate diverse research needs, from cross-linguistic comparisons to multimodal analyses, while ensuring flexibility and interoperability. We detail the modular design of the ontology, its alignment with semantic web standards (RDF/OWL), and the integration of existing ontologies (e.g., OLiA, FOAF). The RCxn’s development is iterative, driven by feedback from our diverse group of PhD researchers.