Stavros Bompolas
2026
Extending ASR Evaluation Resources for Modern Greek Dialects
Chara Tsoukala | Stavros Bompolas | Antigoni Margariti | Konstantina Panagiotou | Maria Elisavet Plaiti | Nefeli Tzanakaki | Petros Karatsareas | Angela Ralli | Antonios Anastasopoulos | Stella Markantonatou
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
Chara Tsoukala | Stavros Bompolas | Antigoni Margariti | Konstantina Panagiotou | Maria Elisavet Plaiti | Nefeli Tzanakaki | Petros Karatsareas | Angela Ralli | Antonios Anastasopoulos | Stella Markantonatou
Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects
Recent progress in Automatic Speech Recognition (ASR) has primarily benefited high-resource standard languages, while dialectal speech remains challenging and underexplored. We present an expanded benchmark for low-resource Modern Greek dialects, covering Aperathiot, Cretan, Lesbian, and Cappadocian, spanning southern, northern, and contact-influenced varieties with varying degrees of divergence from Standard Modern Greek. The benchmark provides dialectal transcriptions in the Greek alphabet, following SMG-based orthographic conventions, while preserving dialectal lexical and morphophonological forms. Using this benchmark, we evaluate state-of-the-art multilingual ASR models in a zero-shot setting and by further fine-tuning per dialect. Zero-shot results reveal a clear performance gradient with dialectal distance from Standard Modern Greek, with best WERs ranging from about 60-70% for southern dialects to over 80% for Lesbian and nearly 97% for Cappadocian. Fine-tuning substantially reduces error rates (up to 47% relative WER improvement), with Cappadocian remaining the most challenging variety (best WER 68.17%). Overall, our results highlight persistent limitations of current pretrained ASR models under dialectal variation and the need for dedicated benchmarks and adaptation strategies.
Structural Divergence under Shared Language-Level Specification: Griko in Universal Dependencies
Stavros Bompolas | Emanuela Pinna | Josep Quer | Marika Lekakou | Stella Markantonatou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Stavros Bompolas | Emanuela Pinna | Josep Quer | Marika Lekakou | Stella Markantonatou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Dialectal varieties pose major challenges for NLP resource development, especially when annotation frameworks are organized around standardized language specifications. In Universal Dependencies (UD), dialects without independent ISO codes are subsumed under the corresponding standard language and inherit its language-level documentation, validator settings, and grammatical inventories. This paper examines Griko, a Greek variety spoken in southern Italy that developed in relative isolation from the Modern Greek dialect continuum while remaining in long-term contact with local Italo-Romance varieties. We assess the consequences of this organizational structure through controlled parsing experiments comparing intra-dialectal training, cross-dialectal transfer from Standard Modern Greek (SMG), script-controlled transfer using romanized SMG, and contact-related cross-lingual transfer from Italian. Our results show that, before romanization, the Italian model even surpasses SMG on several UD metrics and that, although romanization substantially improves SMG-based transfer, performance still remains far below the intra-dialectal baseline. We argue that this persistent gap reflects the interaction between structural divergence and language-level validation constraints, a phenomenon we term ISO-based validation coupling. Through analyses of auxiliary systems, voice marking, and progressive constructions, we show how standard-centric validation architectures can constrain the representation of dialect-specific grammar. More broadly, the Griko case highlights the limitations of language-centric organization in UD and underscores the need for variety-sensitive mechanisms when extending universal annotation frameworks to structurally divergent dialects.
A CLDF-Compliant Lexical Database for Modern Greek Dialects: Resource Design and Dialectometric Analysis
Stavros Bompolas | Natalia Chousou-Polydouri | Manuela Genitsaridi | Danae Karatzanou | Georgios Kostopoulos | Elena Anagnostopoulou | Dimitra Melissaropoulou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Stavros Bompolas | Natalia Chousou-Polydouri | Manuela Genitsaridi | Danae Karatzanou | Georgios Kostopoulos | Elena Anagnostopoulou | Dimitra Melissaropoulou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
This paper presents the first CLDF database that systematically documents lexical variation across 36 Modern Greek varieties (including Standard Modern Greek). The dataset aligns 14,378 lexical items over 345 concepts, links varieties to stable identifiers (Glottocodes), introduces a Greek-specific concept list, and maps meanings to standardized Concepticon concept sets, enabling interoperability and reproducible workflows. To assess whether the database preserves a meaningful dialectological signal, we conduct a dialectometric analysis by computing feature-sensitive string distances over IPA transcriptions and applying hierarchical clustering. The resulting similarity structure recovers major macro-divisions—most notably a broad Northern vs. Southern partition among Koine-descended mainland varieties—and isolates peripheral groups with distinct historical trajectories (e.g., Asia Minor, Italiot, Tsakonian). The database provides scalable infrastructure for quantitative dialectology, comparative Greek linguistics, and dialect-aware language technology.
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Antonis Anastasopoulos | Stella Markantonatou | Angela Ralli | Marcos Zampieri | Stavros Bompolas | Vivian Stamou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
Antonis Anastasopoulos | Stella Markantonatou | Angela Ralli | Marcos Zampieri | Stavros Bompolas | Vivian Stamou
Proceedings of the First Workshop on Dialects in NLP — A Resource Perspective
2025
Crossing Dialectal Boundaries: Building a Treebank for the Dialect of Lesbos through Knowledge Transfer from Standard Modern Greek
Stavros Bompolas | Stella Markantonatou | Angela Ralli | Antonios Anastasopoulos
Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025)
Stavros Bompolas | Stella Markantonatou | Angela Ralli | Antonios Anastasopoulos
Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025)
This paper presents the first treebank for the dialect of Lesbos, a low-resource living Northern variety of Modern Greek (MG), annotated according to the Universal Dependencies (UD) framework. So far, the only dialectal treebank available for Greek developed with cross-dialectal knowledge transfer is an East Cretan one, which belongs to the same Southern branch as Standard Modern Greek (SMG). Our study investigates the effectiveness of cross-dialectal knowledge transfer between dialectologically less similar varieties of the same language by leveraging knowledge from SMG to annotate the Northern dialect of Lesbos. We describe the annotation process, present the resulting treebank, inject additional linguistic knowledge to enhance the results, and evaluate the effectiveness of cross-dialectal knowledge transfer for active annotation. Our findings contribute to a better understanding of how dialectal variation within language families affects knowledge transfer in the UD framework, with implications for other low-resource varieties.
VMWE identification with models trained on GUD (a UDv.2 treebank of Standard Modern Greek)
Stella Markantonatou | Vivian Stamou | Stavros Bompolas | Katerina Anastasopoulou | Irianna Linardaki Vasileiadi | Konstantinos Diamantopoulos | Yannis Kazos | Antonios Anastasopoulos
Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025)
Stella Markantonatou | Vivian Stamou | Stavros Bompolas | Katerina Anastasopoulou | Irianna Linardaki Vasileiadi | Konstantinos Diamantopoulos | Yannis Kazos | Antonios Anastasopoulos
Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025)
UD_Greek-GUD (GUD) is the most recent Universal Dependencies (UD) treebank for Standard Modern Greek (SMG) and the first SMG UD treebank to annotate Verbal Multiword Expressions (VMWEs). GUD contains material from fiction texts and various sites that use colloquial SMG. We describe the special annotation decisions we implemented with GUD, the pipeline we developed to facilitate the active annotation of new material, and we report on the method we designed to evaluate the performance of models trained on GUD as regards VMWE identification tasks.
2018
Search
Fix author
Co-authors
- Stella Markantonatou 5
- Antonios Anastasopoulos 3
- Angela Ralli 3
- Vivian Stamou 2
- Elena Anagnostopoulou 1
- Antonis Anastasopoulos 1
- Katerina Anastasopoulou 1
- Patrizia Belik 1
- Natalia Chousou-Polydouri 1
- Konstantinos Diamantopoulos 1
- Marcello Ferro 1
- Manuela Genitsaridi 1
- Petros Karatsareas 1
- Danae Karatzanou 1
- Yannis Kazos 1
- Georgios Kostopoulos 1
- Marika Lekakou 1
- Antigoni Margariti 1
- Claudia Marzi 1
- Dimitra Melissaropoulou 1
- Ouafae Nahli 1
- Konstantina Panagiotou 1
- Emanuela Pinna 1
- Vito Pirrelli 1
- Maria Elisavet Plaiti 1
- Josep Quer 1
- Chara Tsoukala 1
- Nefeli Tzanakaki 1
- Irianna Linardaki Vasileiadi 1
- Marcos Zampieri 1