Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors

Marvin Kaster, Wei Zhao, Steffen Eger


Abstract
Evaluation metrics are a key ingredient for progress of text generation systems. In recent years, several BERT-based evaluation metrics have been proposed (including BERTScore, MoverScore, BLEURT, etc.) which correlate much better with human assessment of text generation quality than BLEU or ROUGE, invented two decades ago. However, little is known what these metrics, which are based on black-box language model representations, actually capture (it is typically assumed they model semantic similarity). In this work, we use a simple regression based global explainability technique to disentangle metric scores along linguistic factors, including semantics, syntax, morphology, and lexical overlap. We show that the different metrics capture all aspects to some degree, but that they are all substantially sensitive to lexical overlap, just like BLEU and ROUGE. This exposes limitations of these novelly proposed metrics, which we also highlight in an adversarial test scenario.
Anthology ID:
2021.emnlp-main.701
Volume:
Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
Month:
November
Year:
2021
Address:
Online and Punta Cana, Dominican Republic
Editors:
Marie-Francine Moens, Xuanjing Huang, Lucia Specia, Scott Wen-tau Yih
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
8912–8925
Language:
URL:
https://aclanthology.org/2021.emnlp-main.701
DOI:
10.18653/v1/2021.emnlp-main.701
Bibkey:
Cite (ACL):
Marvin Kaster, Wei Zhao, and Steffen Eger. 2021. Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8912–8925, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Cite (Informal):
Global Explainability of BERT-Based Evaluation Metrics by Disentangling along Linguistic Factors (Kaster et al., EMNLP 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.emnlp-main.701.pdf
Video:
 https://aclanthology.org/2021.emnlp-main.701.mp4
Code
 steffeneger/global-explainability-metrics
Data
PAWS