The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems

Daniel Cheng; Kyle Yan; Phillip Keung; Noah A. Smith

The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems

Daniel Cheng, Kyle Yan, Phillip Keung, Noah A. Smith

Abstract

Social media platforms play an increasingly important role as forums for public discourse. Many platforms use recommendation algorithms that funnel users to online groups with the goal of maximizing user engagement, which many commentators have pointed to as a source of polarization and misinformation. Understanding the role of NLP in recommender systems is an interesting research area, given the role that social media has played in world events. However, there are few standardized resources which researchers can use to build models that predict engagement with online groups on social media; each research group constructs datasets from scratch without releasing their version for reuse. In this work, we present a dataset drawn from posts and comments on the online message board Reddit. We develop baseline models for recommending subreddits to users, given the user’s post and comment history. We also study the behavior of our recommender models on subreddits that were banned in June 2020 as part of Reddit’s efforts to stop the dissemination of hate speech.

Anthology ID:: 2022.lrec-1.200
Volume:: Proceedings of the Thirteenth Language Resources and Evaluation Conference
Month:: June
Year:: 2022
Address:: Marseille, France
Editors:: Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, Stelios Piperidis
Venue:: LREC
SIG:
Publisher:: European Language Resources Association
Note:
Pages:: 1885–1889
Language:
URL:: https://aclanthology.org/2022.lrec-1.200/
DOI:
Bibkey:
Cite (ACL):: Daniel Cheng, Kyle Yan, Phillip Keung, and Noah A. Smith. 2022. The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1885–1889, Marseille, France. European Language Resources Association.
Cite (Informal):: The Engage Corpus: A Social Media Dataset for Text-Based Recommender Systems (Cheng et al., LREC 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.lrec-1.200.pdf

PDF Cite Search Fix data