High Accuracy, Low Generalization: Structural Homogeneity and Cross-Dataset Evaluation in Fake-News Benchmarks

Hiram Calvo, Mayte H. Laureano


Abstract
State-of-the-art fake-news classifiers frequently report near-ceiling accuracy on widely used benchmarks such as ISOT, Misinfo, and WELFake. We argue that such results often reflect structural homogeneity and provenance-based separability rather than robust claim-level veracity inference. Anchored in the Information Disorder framework, we analyze how dataset construction operationalizes the notion of “fake” and how this shapes model behavior. We conduct systematic bidirectional cross-dataset experiments across six transfer directions and evaluate performance not only by mean accuracy, but also by variance and directional asymmetry. Results reveal substantial degradation under distribution shift and pronounced transfer asymmetries between dataset pairs. Although not always achieving the highest mean accuracy, affective augmentation combining dimensional (VAD) and categorical (Ekman) representations yields the lowest variance and smallest directional gap, indicating superior cross-domain stability. Our findings expose the disconnect between accuracy-driven benchmarking and construct-valid evaluation. We argue that progress in fake-news detection requires shifting from isolated in-domain optimization toward robustness-oriented, bidirectional, and distribution-aware assessment practices.
Anthology ID:
2026.indor-1.3
Volume:
Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Simona Frenda, Marco Antonio Stranisci, Shaina Ashraf, Ada Ren, Ioannis Konstas, Usman Naseem
Venues:
InDor | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
25–33
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-indor-03
DOI:
10.63317/4ka6fv4hbw9z
Bibkey:
Cite (ACL):
Hiram Calvo and Mayte H. Laureano. 2026. High Accuracy, Low Generalization: Structural Homogeneity and Cross-Dataset Evaluation in Fake-News Benchmarks. In Proceedings of the 1st Workshop on Information Disorder (InDor) @ LREC 2026, pages 25–33, Palma de Mallorca, Spain. ELRA Language Resources Association (ELRA).
Cite (Informal):
High Accuracy, Low Generalization: Structural Homogeneity and Cross-Dataset Evaluation in Fake-News Benchmarks (Calvo & Laureano, InDor 2026)
Copy Citation: