JobResQA: Semi-Automatic Multilingual Benchmark Creation for LLM Machine Reading Comprehension on Résumés and Job Descriptions

Casimiro Pio Carrino, Paula Estrella, Rabih Zbib, Carlos Escolano, Jose A. R. Fonollosa


Abstract
We present a methodology for building privacy-preserving multilingual QA benchmarks in low-resource and sensitive domains, demonstrated through JobResQA, a multilingual MRC benchmark over synthetic HR documents. The dataset comprises 581 QA pairs across 105 synthetic résumé-job description pairs in five languages (English, Spanish, Italian, German, and Chinese), with questions spanning four types based on document source (intra vs. cross-document) and reasoning complexity (single-hop vs. multi-hop). We propose a privacy-preserving synthetic data pipeline applicable to other sensitive domains, with controlled demographic attributes (via placeholders) enabling future bias studies. Our cost-effective, human-in-the-loop translation pipeline based on TEaR methodology incorporates MQM error annotations and selective post-editing. Baseline evaluations across multiple open-weight LLM families using LLM-as-judge reveal higher performance on English and Spanish but substantial degradation for other languages, highlighting critical cross-lingual MRC gaps. Our pipeline, where LLMs act as synthesizers, translators, and evaluators under human oversight, constitutes a reusable methodology for resource creation and a case study in evaluation-integrity challenges of LLM-era benchmark construction.
Anthology ID:
2026.resourceful-4.15
Volume:
Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026)
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Felix Morger, Nikolai Ilinykh, Barbara Scalvini, Simon Dobnik, Dana Dannélls
Venues:
RESOURCEFUL | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
161–176
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-resourceful-15
DOI:
10.63317/4pzwmqxt46xp
Bibkey:
Cite (ACL):
Casimiro Pio Carrino, Paula Estrella, Rabih Zbib, Carlos Escolano, and Jose A. R. Fonollosa. 2026. JobResQA: Semi-Automatic Multilingual Benchmark Creation for LLM Machine Reading Comprehension on Résumés and Job Descriptions. In Proceedings of the Fourth Workshop on the Role of Resources in the Age of Large Language Models (RESOURCEFUL 2026), pages 161–176, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
JobResQA: Semi-Automatic Multilingual Benchmark Creation for LLM Machine Reading Comprehension on Résumés and Job Descriptions (Carrino et al., RESOURCEFUL 2026)
Copy Citation: