CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking

Xuming Hu; Zhijiang Guo; GuanYu Wu; Aiwei Liu; Lijie Wen; Philip S. Yu

doi:10.18653/v1/2022.naacl-main.246

CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking

Xuming Hu, Zhijiang Guo, GuanYu Wu, Aiwei Liu, Lijie Wen, Philip Yu

Abstract

The explosion of misinformation spreading in the media ecosystem urges for automated fact-checking. While misinformation spans both geographic and linguistic boundaries, most work in the field has focused on English. Datasets and tools available in other languages, such as Chinese, are limited. In order to bridge this gap, we construct CHEF, the first CHinese Evidence-based Fact-checking dataset of 10K real-world claims. The dataset covers multiple domains, ranging from politics to public health, and provides annotated evidence retrieved from the Internet. Further, we develop established baselines and a novel approach that is able to model the evidence retrieval as a latent variable, allowing jointly training with the veracity prediction model in an end-to-end fashion. Extensive experiments show that CHEF will provide a challenging testbed for the development of fact-checking systems designed to retrieve and reason over non-English claims.

Anthology ID:: 2022.naacl-main.246
Volume:: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Month:: July
Year:: 2022
Address:: Seattle, United States
Editors:: Marine Carpuat, Marie-Catherine de Marneffe, Ivan Vladimir Meza Ruiz
Venue:: NAACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 3362–3376
Language:
URL:: https://aclanthology.org/2022.naacl-main.246/
DOI:: 10.18653/v1/2022.naacl-main.246
Bibkey:
Cite (ACL):: Xuming Hu, Zhijiang Guo, GuanYu Wu, Aiwei Liu, Lijie Wen, and Philip Yu. 2022. CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3362–3376, Seattle, United States. Association for Computational Linguistics.
Cite (Informal):: CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking (Hu et al., NAACL 2022)
Copy Citation:
PDF:: https://aclanthology.org/2022.naacl-main.246.pdf
Video:: https://aclanthology.org/2022.naacl-main.246.mp4

PDF Cite Search Video Fix data