Oliver Pejić
2026
Thematic Landscapes of the Past: Analysing Slovene Historical Periodicals With Topic Modeling
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Tina Munda | Darja Fiser
Proceedings of the First Workshop on Creating Interoperable Corpora of Historical Newspapers
This paper explores the thematic landscapes of three Slovene historical periodicals—Slovenka, Slovenec, and Slovenski narod—from the sPeriodika corpus, a comprehensive collection of Slovene press published between 1771 and 1914. Using BERTopic, we analyse the thematic profiles of these periodicals, enriched with diachronic perspectives. Our study examines the thematic commonalities and specificities of the selected periodicals, highlighting their distinct political orientations, target audiences, and the increasing nationalist polarisation in public discourse. This work contributes to digital humanities by demonstrating the potential of modern topic modelling techniques, such as BERTopic, to advance historical and cultural research.
Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Tina Munda | Filip Dobranić | Uroš Šmajdek | Oliver Pejić | Ciril Bohak | Vojko Gorjanc | Darja Fišer
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Historical newspapers present substantial challenges for computational sentiment analysis due to OCR noise, archaic linguistic features, and the absence of domain-specific labeled training data. This paper examines whether instruction-following LLMs can support targeted, mention-level sentiment inference in such conditions. We benchmark four instruction-following LLMs on a manually annotated sample of collective-identity mentions drawn from Slovene historical newspapers. The results provide a benchmark for targeted sentiment classification in OCR-degraded historical Slovene and offer an empirically grounded assessment of the capabilities and limitations of an instruction-tuned LLM in digital humanities research.