Investigating African-American Vernacular English in Transformer-Based Text Generation

Sophie Groenwold; Lily Ou; Aesha Parekh; Samhita Honnavalli; Sharon Levy; Diba Mirza; William Yang Wang

doi:10.18653/v1/2020.emnlp-main.473

Investigating African-American Vernacular English in Transformer-Based Text Generation

Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, William Yang Wang

Abstract

The growth of social media has encouraged the written use of African American Vernacular English (AAVE), which has traditionally been used only in oral contexts. However, NLP models have historically been developed using dominant English varieties, such as Standard American English (SAE), due to text corpora availability. We investigate the performance of GPT-2 on AAVE text by creating a dataset of intent-equivalent parallel AAVE/SAE tweet pairs, thereby isolating syntactic structure and AAVE- or SAE-specific language for each pair. We evaluate each sample and its GPT-2 generated text with pretrained sentiment classifiers and find that while AAVE text results in more classifications of negative sentiment than SAE, the use of GPT-2 generally increases occurrences of positive sentiment for both. Additionally, we conduct human evaluation of AAVE and SAE text generated with GPT-2 to compare contextual rigor and overall quality.

Anthology ID:: 2020.emnlp-main.473
Volume:: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
Month:: November
Year:: 2020
Address:: Online
Editors:: Bonnie Webber, Trevor Cohn, Yulan He, Yang Liu
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 5877–5883
Language:
URL:: https://aclanthology.org/2020.emnlp-main.473
DOI:: 10.18653/v1/2020.emnlp-main.473
Bibkey:
Cite (ACL):: Sophie Groenwold, Lily Ou, Aesha Parekh, Samhita Honnavalli, Sharon Levy, Diba Mirza, and William Yang Wang. 2020. Investigating African-American Vernacular English in Transformer-Based Text Generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5877–5883, Online. Association for Computational Linguistics.
Cite (Informal):: Investigating African-American Vernacular English in Transformer-Based Text Generation (Groenwold et al., EMNLP 2020)
Copy Citation:
PDF:: https://aclanthology.org/2020.emnlp-main.473.pdf
Optional supplementary material:: 2020.emnlp-main.473.OptionalSupplementaryMaterial.zip
Video:: https://slideslive.com/38939366
Code: sophiegroenwold/AAVE_SAE_dataset
Data: AAVE/SAE Paired Dataset

PDF Cite Search Code Optional supplementary material Video