Re-using an Argument Corpus to Aid in the Curation of Social Media Collections

Clare Llewellyn; Claire Grover; Jon Oberlander; Ewan Klein

Re-using an Argument Corpus to Aid in the Curation of Social Media Collections

Clare Llewellyn, Claire Grover, Jon Oberlander, Ewan Klein

Abstract

This work investigates how automated methods can be used to classify social media text into argumentation types. In particular it is shown how supervised machine learning was used to annotate a Twitter dataset (London Riots) with argumentation classes. An investigation of issues arising from a natural inconsistency within social media data found that machine learning algorithms tend to over fit to the data because Twitter contains a lot of repetition in the form of retweets. It is also noted that when learning argumentation classes we must be aware that the classes will most likely be of very different sizes and this must be kept in mind when analysing the results. Encouraging results were found in adapting a model from one domain of Twitter data (London Riots) to another (OR2012). When adapting a model to another dataset the most useful feature was punctuation. It is probable that the nature of punctuation in Twitter language, the very specific use in links, indicates argumentation class.

Anthology ID:: L14-1651
Volume:: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)
Month:: May
Year:: 2014
Address:: Reykjavik, Iceland
Editors:: Nicoletta Calzolari, Khalid Choukri, Thierry Declerck, Hrafn Loftsson, Bente Maegaard, Joseph Mariani, Asuncion Moreno, Jan Odijk, Stelios Piperidis
Venue:: LREC
SIG:
Publisher:: European Language Resources Association (ELRA)
Note:
Pages:: 462–468
Language:
URL:: http://www.lrec-conf.org/proceedings/lrec2014/pdf/845_Paper.pdf
DOI:
Bibkey:
Cite (ACL):: Clare Llewellyn, Claire Grover, Jon Oberlander, and Ewan Klein. 2014. Re-using an Argument Corpus to Aid in the Curation of Social Media Collections. In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14), pages 462–468, Reykjavik, Iceland. European Language Resources Association (ELRA).
Cite (Informal):: Re-using an Argument Corpus to Aid in the Curation of Social Media Collections (Llewellyn et al., LREC 2014)
Copy Citation:
PDF:: http://www.lrec-conf.org/proceedings/lrec2014/pdf/845_Paper.pdf

PDF Cite Search Fix data