Haohong Lai

Author directory

2026

This study examines the influence of prompt language and translation theory-driven prompt design on the quality of Spanish–Chinese editorial translations generated by GPT-5.2. A parallel corpus of four EL PAÍS editorials was translated under 48 experimental conditions (4 prompt types × 3 prompt languages × 4 articles). Translation quality was assessed using BLEU and BERTScore-F1 for automated evaluation, alongside human evaluation based on the Multidimensional Quality Metrics (MQM) framework. Automated metrics identified the baseline prompt (BASE) as the best-performing condition, whereas human evaluation ranked the brief-oriented prompt (BRIEF) highest (MQM: 8.66 vs. 7.84), a reversal attributed to the single-reference constraint inherent in automated measures. Subtype analysis indicated that translation theory-driven prompts selectively reduced Awkward style errors, whereas Unidiomatic style errors persisted consistently across conditions. Prompt language exhibited negligible impact under both evaluation paradigms. These results indicate that translation theory-driven prompts are advantageous for language learners seeking high-quality editorial translations and underscore the necessity of human evaluation for accurately assessing LLM translation quality in this domain.