Zero-shot Cross-lingual POS Tagging for Filipino

Jimson Paulo Layacan; Isaiah Edri W. Flores; Katrina Bernice M. Tan; Ma. Regina E. Estuar; Jann Railey Montalan; Marlene M. De Leon

doi:10.18653/v1/2024.fieldmatters-1.9

Zero-shot Cross-lingual POS Tagging for Filipino

Jimson Paulo Layacan, Isaiah Edri W. Flores, Katrina Bernice M. Tan, Ma. Regina E. Estuar, Jann Railey E. Montalan, Marlene M. De Leon

Abstract

Supervised learning approaches in NLP, exemplified by POS tagging, rely heavily on the presence of large amounts of annotated data. However, acquiring such data often requires significant amount of resources and incurs high costs. In this work, we explore zero-shot cross-lingual transfer learning to address data scarcity issues in Filipino POS tagging, particularly focusing on optimizing source language selection. Our zero-shot approach demonstrates superior performance compared to previous studies, with top-performing fine-tuned PLMs achieving F1 scores as high as 79.10%. The analysis reveals moderate correlations between cross-lingual transfer performance and specific linguistic distances–featural, inventory, and syntactic–suggesting that source languages with these features closer to Filipino provide better results. We identify tokenizer optimization as a key challenge, as PLM tokenization sometimes fails to align with meaningful representations, thus hindering POS tagging performance.

Anthology ID:: 2024.fieldmatters-1.9
Volume:: Proceedings of the Third Workshop on NLP Applications to Field Linguistics
Month:: August
Year:: 2024
Address:: Bangkok, Thailand
Editors:: Oleg Serikov, Ekaterina Voloshina, Anna Postnikova, Saliha Muradoglu, Eric Le Ferrand, Elena Klyachko, Ekaterina Vylomova, Tatiana Shavrina, Francis Tyers
Venues:: FieldMatters | WS
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 69–77
Language:
URL:: https://aclanthology.org/2024.fieldmatters-1.9/
DOI:: 10.18653/v1/2024.fieldmatters-1.9
Bibkey:
Cite (ACL):: Jimson Paulo Layacan, Isaiah Edri W. Flores, Katrina Bernice M. Tan, Ma. Regina E. Estuar, Jann Railey E. Montalan, and Marlene M. De Leon. 2024. Zero-shot Cross-lingual POS Tagging for Filipino. In Proceedings of the Third Workshop on NLP Applications to Field Linguistics, pages 69–77, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):: Zero-shot Cross-lingual POS Tagging for Filipino (Layacan et al., FieldMatters 2024)
Copy Citation:
PDF:: https://aclanthology.org/2024.fieldmatters-1.9.pdf

PDF Cite Search Fix data