Modeling the "Dalet" Clitic in Historical Hebrew Texts: A New Prefix-Segmented BERT Model and Stylistic Analysis

Rachel Tal; Cheyn Shmuel Shmidman; Avi Shmidman

Modeling the "Dalet" Clitic in Historical Hebrew Texts: A New Prefix-Segmented BERT Model and Stylistic Analysis

Rachel Tal, Cheyn Shmuel Shmidman, Avi Shmidman

Abstract

The Aramaic proclitic *dalet*, widely used in historical Hebrew texts, serves two distinct grammatical functions: as a subordinating conjunction and as a possessive preposition. Because these functions are orthographically identical and no annotated resources exist for this task, large-scale computational analysis of their usage has previously been infeasible. In this paper we introduce a new BERT model for historical Hebrew in which all prefixes are segmented and encoded as independent tokens. This representation allows the model to evaluate proclitics directly and provides a probe-based unsupervised method for determining the grammatical role of the *dalet* clitic using masked language modeling predictions. We evaluate the approach on a manually annotated dataset drawn from historical Hebrew literature spanning multiple regions and historical periods, achieving over an average F1 score of over 0.89. Applying the method to a corpus of more than 300 million words of historical Hebrew texts, we conduct large-scale stylistic analyses of the choice between the Aramaic *dalet* and available Hebrew alternatives. The results reveal geographic and diachronic trends and identify distinct stylistic clusters within the corpus. The prefix-segmented model and annotated dataset are released for unrestricted use.

Anthology ID:: 2026.nlp4dh-1.12
Volume:: Proceedings of the 6th International Conference on Natural Language Processing for the Digital Humanities
Month:: July
Year:: 2026
Address:: San Diego, USA
Editors:: Sil Hamilton, Emily Öhman, Rebecca M. M. Hicke, Yuri Bizzoni, Axel Bax, Jacob A. Matthews, Mika Hämäläinen
Venues:: NLP4DH | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 121–131
Language:
URL:: https://aclanthology.org/2026.nlp4dh-1.12/
DOI:
Bibkey:
Cite (ACL):: Rachel Tal, Cheyn Shmuel Shmidman, and Avi Shmidman. 2026. Modeling the "Dalet" Clitic in Historical Hebrew Texts: A New Prefix-Segmented BERT Model and Stylistic Analysis. In Proceedings of the 6th International Conference on Natural Language Processing for the Digital Humanities, pages 121–131, San Diego, USA. Association for Computational Linguistics.
Cite (Informal):: Modeling the “Dalet” Clitic in Historical Hebrew Texts: A New Prefix-Segmented BERT Model and Stylistic Analysis (Tal et al., NLP4DH 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.nlp4dh-1.12.pdf

PDF Cite Search Fix data