%0 Conference Proceedings %T Towards Actual (Not Operational) Textual Style Transfer Auto-Evaluation %A Pang, Richard Yuanzhe %Y Xu, Wei %Y Ritter, Alan %Y Baldwin, Tim %Y Rahimi, Afshin %S Proceedings of the 5th Workshop on Noisy User-generated Text (W-NUT 2019) %D 2019 %8 November %I Association for Computational Linguistics %C Hong Kong, China %F pang-2019-towards %X Regarding the problem of automatically generating paraphrases with modified styles or attributes, the difficulty lies in the lack of parallel corpora. Numerous advances have been proposed for the generation. However, significant problems remain with the auto-evaluation of style transfer tasks. Based on the summary of Pang and Gimpel (2018) and Mir et al. (2019), style transfer evaluations rely on three metrics: post-transfer style classification accuracy, content or semantic similarity, and naturalness or fluency. We elucidate the dangerous current state of style transfer auto-evaluation research. Moreover, we propose ways to aggregate the three metrics into one evaluator. This abstract aims to bring researchers to think about the future of style transfer and style transfer evaluation research. %R 10.18653/v1/D19-5557 %U https://aclanthology.org/D19-5557 %U https://doi.org/10.18653/v1/D19-5557 %P 444-445