Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text

Louis Estève, Agata Savary, Thomas Lavergne


Abstract
Multiword Expressions (MWEs) make a goodcase study for linguistic diversity due to theiridiosyncratic nature. Defining MWE canonicalforms as types, diversity may be measurednotably through disparity, based on pairwisedistances between types. To this aim, wetrain static MWE-aware word embeddings forverbal MWEs in 14 languages, and we showinteresting properties of these vector spaces.We use these vector spaces to implement theso-called functional diversity measure. Weapply this measure to the results of severalMWE identification systems. We find that,although MWE vector spaces are meaningful ata local scale, the disparity measure aggregatingthem at a global scale strongly correlateswith the number of types, which questions itsusefulness in presence of simpler diversitymetrics such as variety. We make the vectorspaces we generated available.
Anthology ID:
2024.acl-srw.20
Volume:
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)
Month:
August
Year:
2024
Address:
Bangkok, Thailand
Editors:
Xiyan Fu, Eve Fleisig
Venue:
ACL
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
110–130
Language:
URL:
https://aclanthology.org/2024.acl-srw.20
DOI:
10.18653/v1/2024.acl-srw.20
Bibkey:
Cite (ACL):
Louis Estève, Agata Savary, and Thomas Lavergne. 2024. Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 110–130, Bangkok, Thailand. Association for Computational Linguistics.
Cite (Informal):
Vector Spaces for Quantifying Disparity of Multiword Expressions in Annotated Text (Estève et al., ACL 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.acl-srw.20.pdf