Taras Andrushko - ACL Anthology

Taras Andrushko

2022

UniMorph 4.0: Universal Morphology
Khuyagbaatar Batsuren | Omer Goldman | Salam Khalifa | Nizar Habash | Witold Kieraś | Gábor Bella | Brian Leonard | Garrett Nicolai | Kyle Gorman | Yustinus Ghanggo Ate | Maria Ryskina | Sabrina Mielke | Elena Budianskaya | Charbel El-Khaissi | Tiago Pimentel | Michael Gasser | William Abbott Lane | Mohit Raj | Matt Coler | Jaime Rafael Montoya Samame | Delio Siticonatzi Camaiteri | Esaú Zumaeta Rojas | Didier López Francis | Arturo Oncevay | Juan López Bautista | Gema Celeste Silva Villegas | Lucas Torroba Hennigen | Adam Ek | David Guriel | Peter Dirix | Jean-Philippe Bernardy | Andrey Scherbakov | Aziyana Bayyr-ool | Antonios Anastasopoulos | Roberto Zariquiey | Karina Sheifer | Sofya Ganieva | Hilaria Cruz | Ritván Karahóǧa | Stella Markantonatou | George Pavlidis | Matvey Plugaryov | Elena Klyachko | Ali Salehi | Candy Angulo | Jatayu Baxi | Andrew Krizhanovsky | Natalia Krizhanovskaya | Elizabeth Salesky | Clara Vania | Sardana Ivanova | Jennifer White | Rowan Hall Maudslay | Josef Valvoda | Ran Zmigrod | Paula Czarnowska | Irene Nikkarinen | Aelita Salchak | Brijesh Bhatt | Christopher Straughn | Zoey Liu | Jonathan North Washington | Yuval Pinter | Duygu Ataman | Marcin Wolinski | Totok Suhardijanto | Anna Yablonskaya | Niklas Stoehr | Hossep Dolatian | Zahroh Nuriah | Shyam Ratan | Francis M. Tyers | Edoardo M. Ponti | Grant Aiton | Aryaman Arora | Richard J. Hatcher | Ritesh Kumar | Jeremiah Young | Daria Rodionova | Anastasia Yemelina | Taras Andrushko | Igor Marchenko | Polina Mashkovtseva | Alexandra Serova | Emily Prud’hommeaux | Maria Nepomniashchaya | Fausto Giunchiglia | Eleanor Chodroff | Mans Hulden | Miikka Silfverberg | Arya D. McCarthy | David Yarowsky | Ryan Cotterell | Reut Tsarfaty | Ekaterina Vylomova
Proceedings of the Thirteenth Language Resources and Evaluation Conference

The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation, and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements on several fronts that were made in the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 66 new languages, including 24 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g., missing gender and macrons information. We have amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet.

The 2022 SIGMORPHON–UniMorph shared task on large scale morphological inflection generation included a wide range of typologically diverse languages: 33 languages from 11 top-level language families: Arabic (Modern Standard), Assamese, Braj, Chukchi, Eastern Armenian, Evenki, Georgian, Gothic, Gujarati, Hebrew, Hungarian, Itelmen, Karelian, Kazakh, Ket, Khalkha Mongolian, Kholosi, Korean, Lamahalot, Low German, Ludic, Magahi, Middle Low German, Old English, Old High German, Old Norse, Polish, Pomak, Slovak, Turkish, Upper Sorbian, Veps, and Xibe. We emphasize generalization along different dimensions this year by evaluating test items with unseen lemmas and unseen features separately under small and large training conditions. Across the five submitted systems and two baselines, the prediction of inflections with unseen features proved challenging, with average performance decreased substantially from last year. This was true even for languages for which the forms were in principle predictable, which suggests that further work is needed in designing systems that capture the various types of generalization required for the world’s languages.

Co-authors

Elena Budianskaya 2

Ryan Cotterell 2

Hossep Dolatian 2

Salam Khalifa 2

Witold Kieraś 2

Andrew Krizhanovsky 2

Igor Marchenko 2

Polina Mashkovtseva 2

Maria Nepomniashchaya 2

Daria Rodionova 2

Ekaterina Vylomova 2

Anastasia Yemelina 2

Jeremiah Young 2

Nona Atanalov 1

Aziyana Bayyr-ool 1

Jean-Philippe Bernardy 1

Brijesh Bhatt 1

Delio Siticonatzi Camaiteri 1

Eleanor Chodroff 1

Paula Czarnowska 1

Charbel El-Khaissi 1

Sofya Ganieva 1

Michael Gasser 1

Fausto Giunchiglia 1

Silvia Guriel-Agiashvili 1

Richard J. Hatcher 1

Sardana Ivanova 1

Ritván Karahóǧa 1

Elena Klyachko 1

Jordan Kodner 1

Natalia Krizhanovskaya 1

Natalia Krizhanovsky 1

William Abbott Lane 1

Brian Leonard 1

Juan López Bautista 1

Didier López Francis 1

Stella Markantonatou 1

Magdalena Markowska 1

Rowan Hall Maudslay 1

Arya D. McCarthy 1

Sabrina J. Mielke 1

Garrett Nicolai 1

Irene Nikkarinen 1

Zahroh Nuriah 1

Arturo Oncevay 1

George Pavlidis 1

Tiago Pimentel 1

Matvey Plugaryov 1

Edoardo M. Ponti 1

Emily Prud’hommeaux 1

Esaú Zumaeta Rojas 1

Maria Ryskina 1

Aelita Salchak 1

Elizabeth Salesky 1

Jaime Rafael Montoya Samame 1

Karina Scheifer 1

Andrey Scherbakov 1

Alexandra Serova 1

Karina Sheifer 1

Miikka Silfverberg 1

Alexandra Sorova 1

Niklas Stoehr 1

Christopher Straughn 1

Totok Suhardijanto 1

Lucas Torroba Hennigen 1

Reut Tsarfaty 1

Francis Tyers 1

Josef Valvoda 1

Gema Celeste Silva Villegas 1

Jonathan Washington 1

Jennifer White 1

Marcin Woliński 1

Anna Yablonskaya 1

David Yarowsky 1

Roberto Zariquiey 1

Venues

LREC1
SIGMORPHON1