Across Generations: A Comparative Analysis of NER for Latin Inscriptions from Classical Machine Learning to LLMs

Wenhui Cui, Phillip Benjamin Ströbel


Abstract
Latin epigraphic texts are a challenging type of historical data for natural language processing (NLP). They are often fragmentary, contain inconsistent spelling, and follow complex Roman naming conventions. This paper investigates Named Entity Recognition (NER) for this domain by comparing several approaches, including feature-based Support Vector Machines, neural models such as BiLSTM and TreeLSTM, pre-trained language models like LatinBERT, fine-tuned Transformer models based on BERT, and large language models used with prompting and supervised fine-tuning. We introduce a manually annotated dataset of 1,000 inscriptions from the Epigraphik-Datenbank Clauss-Slaby, labelled with a fine-grained BIO scheme that captures the internal structure of Roman personal names. Results show that the fine-tuned BERT model achieves the highest performance, with a weighted F1 score of 91.1% and a macro F1 of 68.7%, and clearly outperforms other methods. Additional linguistic features, such as part-of-speech tags and dependency information, yield only limited improvements, likely due to the irregular nature of inscriptional texts. This work provides a new benchmark for NER on Latin inscriptions and offers practical insights into applying modern NLP techniques to historical, non-standardised language.
Anthology ID:
2026.lt4hala-1.11
Volume:
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Rachele Sprugnoli, Marco Passarotti
Venues:
LT4HALA | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
112–124
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-lt4hala-11
DOI:
10.63317/2g99sovd35pj
Bibkey:
Cite (ACL):
Wenhui Cui and Phillip Benjamin Ströbel. 2026. Across Generations: A Comparative Analysis of NER for Latin Inscriptions from Classical Machine Learning to LLMs. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026, pages 112–124, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Across Generations: A Comparative Analysis of NER for Latin Inscriptions from Classical Machine Learning to LLMs (Cui & Ströbel, LT4HALA 2026)
Copy Citation: