Robustness Evaluation of Text Classification Models Using Mathematical Optimization and Its Application to Adversarial Training

Hikaru Tomonari; Masaaki Nishino; Akihiro Yamamoto

doi:10.18653/v1/2022.findings-aacl.31

Robustness Evaluation of Text Classification Models Using Mathematical Optimization and Its Application to Adversarial Training

Hikaru Tomonari, Masaaki Nishino, Akihiro Yamamoto

Abstract

Neural networks are known to be vulnerable to adversarial examples due to slightly perturbed input data. In practical applications of neural network models, the robustness of the models against perturbations must be evaluated. However, no method can strictly evaluate their robustness in natural language domains. We therefore propose a method that evaluates the robustness of text classification models using an integer linear programming (ILP) solver by an optimization problem that identifies a minimum synonym swap that changes the classification result. Our method allows us to compare the robustness of various models in realistic time. It can also be used for obtaining adversarial examples. Because of the minimal impact on the altered sentences, adversarial examples with our method obtained high scores in human evaluations of grammatical correctness and semantic similarity for an IMDb dataset. In addition, we implemented adversarial training with the IMDb and SST2 datasets and found that our adversarial training method makes the model robust.

Anthology ID:: 2022.findings-aacl.31
Volume:: Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022
Month:: November
Year:: 2022
Address:: Online only
Editors:: Yulan He, Heng Ji, Sujian Li, Yang Liu, Chua-Hui Chang
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 327–333
Language:
URL:: https://aclanthology.org/2022.findings-aacl.31/
DOI:: 10.18653/v1/2022.findings-aacl.31
Bibkey:
Cite (ACL):: Hikaru Tomonari, Masaaki Nishino, and Akihiro Yamamoto. 2022. Robustness Evaluation of Text Classification Models Using Mathematical Optimization and Its Application to Adversarial Training. In Findings of the Association for Computational Linguistics: AACL-IJCNLP 2022, pages 327–333, Online only. Association for Computational Linguistics.
Cite (Informal):: Robustness Evaluation of Text Classification Models Using Mathematical Optimization and Its Application to Adversarial Training (Tomonari et al., Findings 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.findings-aacl.31.pdf

PDF Cite Search Fix data