M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

Hao Wang; Linlong Xu; Heng Liu; Yangyang Liu; Xiaohu Zhao; Bo Zeng; Liangying Shao; Yichen Dong; Xinwei Wu; Jiang Zhou; Tianyu Dong; Xiangxiang Zeng; Longyue Wang; Weihua Luo

M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation

Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo

Abstract

Aligning Large Language Models (LLMs) to human preferences is pivotal for Machine Translation (MT), yet current approaches are often hindered by misleading reward signals. Our analysis reveals that prevailing Quality Estimation (QE) models exhibit a systematic blind spot towards **partial errors**—specifically partial hallucinations and omissions—often favoring superficially fluent but unfaithful translations. To address this, we propose **M²PO** (**M**ulti-Perspective **M**ulti-Pair **P**reference **O**ptimization), a data-centric framework for preference optimization in machine translation. First, to correct the bias towards fluency, M²PO uses a multi-perspective alignment mechanism that decouples semantic fidelity from fluency, prioritizing faithfulness via a curriculum strategy. Second, with the bias corrected, partial errors fall between perfect and severely incorrect translations, making them inefficient to learn via standard best-versus-worst comparisons. We thus introduce a multi-pair objective that leverages the full candidate list to capture these fine-grained error signals. Experiments on WMT23, WMT24, and FLORES-200 show that M²PO enables a 9B model to outperform leading open-source baselines and achieve parity with proprietary models like GPT-4o and Gemini-2.0-Flash, demonstrating significant potential for efficient, high-fidelity LLM-based translation.

Anthology ID:: 2026.acl-long.469
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 10315–10336
Language:
URL:: https://aclanthology.org/2026.acl-long.469/
DOI:
Bibkey:
Cite (ACL):: Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, and Weihua Luo. 2026. M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10315–10336, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: M2PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation (Wang et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.469.pdf
Checklist:: 2026.acl-long.469.checklist.pdf

PDF Cite Search Checklist Fix data