Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals

Tina Munda, Filip Dobranić, Uroš Šmajdek, Oliver Pejić, Ciril Bohak, Vojko Gorjanc, Darja Fišer


Abstract
Historical newspapers present substantial challenges for computational sentiment analysis due to OCR noise, archaic linguistic features, and the absence of domain-specific labeled training data. This paper examines whether instruction-following LLMs can support targeted, mention-level sentiment inference in such conditions. We benchmark four instruction-following LLMs on a manually annotated sample of collective-identity mentions drawn from Slovene historical newspapers. The results provide a benchmark for targeted sentiment classification in OCR-degraded historical Slovene and offer an empirically grounded assessment of the capabilities and limitations of an instruction-tuned LLM in digital humanities research.
Anthology ID:
2026.llms4ssh-1.11
Volume:
Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca (Spain)
Editors:
Arturo Montejo-Raez, Cristina Grisot, Joanna Blochowiak, Nikola Ljubešić, Elena Battaner, German Rigau
Venues:
LLMs4SSH | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
103–113
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-llms4ssh-11
DOI:
10.63317/22hvcbc23rts
Bibkey:
Cite (ACL):
Tina Munda, Filip Dobranić, Uroš Šmajdek, Oliver Pejić, Ciril Bohak, Vojko Gorjanc, and Darja Fišer. 2026. Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals. In Proceedings of Shaping Multilingual, Multimodal AI for the Social Sciences and Humanities (LLMs4SSH) @ LREC 2026, pages 103–113, Palma de Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Benchmarking LLMs for Aspect-Based Sentiment Classification in Slovene Historical Periodicals (Munda et al., LLMs4SSH 2026)
Copy Citation: