Lisa Sophie Albertelli
2026
Modeling Topics as Linguistic Linked Open Data: A First Attempt Using BERTopic, Ontolex-Lemon and FrAC
Lisa Sophie Albertelli
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
Lisa Sophie Albertelli
Proceedings of 10th Workshop on Linked Data in Linguistics (LDL-2026)
Parliamentary discourse constitutes a key domain in which political actors publicly articulate policy positions and priorities through language. This study investigates debates from the Italian Chamber of Deputies (1948–2006) to identify and analyse latent semantic themes and their evolution using BERTopic-based dynamic topic modeling. The analysis relies on a subset of the ItaParlCorpus (Cova, 2025), a large-scale, machine-readable corpus enriched with temporal, institutional, and political metadata. Beyond topic extraction,this work addresses a largely unexplored challenge: the formalization of topics derived from unsupervised, embedding-based topic modeling as Linked Data entities, adopting a linguistic perspective. Extracted topics are formalized as semantic entities reusing the OntoLex–Lemon model, its FrAC extension and declaring a dedicated ontology to link topics to speeches, speakers, political parties, and temporal information reusing standardized vocabularies and persistent URIs. This integration enables semantic querying through SPARQL, supporting analyses of topic distributions across political actors, parties and illustrating the analytical potential of the proposed approach. Moreover, the study highlights limitations in the formalization of topic modeling outputs, particularly regarding the representation of ambiguous word forms and their alignment with lexical concepts in OntoLex–Lemon.
2025
The Leibniz List as Linguistic Linked Data in the LiLa Knowledge Base
Lisa Sophie Albertelli | Giulia Calvi | Francesco Mambrini
Proceedings of the 5th Conference on Language, Data and Knowledge
Lisa Sophie Albertelli | Giulia Calvi | Francesco Mambrini
Proceedings of the 5th Conference on Language, Data and Knowledge
This paper presents the integration of the Leibniz List, a concept list from the Concepticon project, into the LiLa Knowledge Base of Latin interoperable resources. The modeling experiment was conducted using W3C standards like Ontolex and SKOS. This work, which originated in a project for a university course, is limited to a short list of words, but it already enables interoperability between the Concepticon and the language resources in a LOD architecture like LiLa. The integration enriches the LiLa ecosystem, allowing users to explore Latin lexicon from an onomasiological perspective and links concepts to lexical entries from various dictionaries and corpus attestations. The work showcases how standard Semantic Web technologies can effectively model and connect historical concept lists within larger linguistic knowledge infrastructures and provides an example for further experiments with the Concepticon’s data.