Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

Jian Li; Shenglin Yin; Yujia Zhang; Alan Zhao; Xi Chen; Xiaohui Zhou; Pengfei Xu

doi:10.18653/v1/2025.emnlp-main.460

Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization

Jian Li, Shenglin Yin, Yujia Zhang, Alan Zhao, Xi Chen, Xiaohui Zhou, Pengfei Xu

Abstract

Direct Preference Optimization (DPO) is a widely used reinforcement learning from human feedback (RLHF) method across various domains. The study of token importance has attracted widespread attention in DPO. Researchers have found that token importance is crucial for improving the effectiveness of DPO. It is observed that identical or semantically similar content (defined as ambiguous content) frequently appears within the preference pairs. We hypothesize that the presence of ambiguous content during DPO training may introduce ambiguity, thereby limiting further improvements in alignment. Through mathematical analysis and proof-of-concept experiments, we reveal that ambiguous content may potentially introduce ambiguities, thereby degrading performance. To address this issue, we introduce Ambiguity Awareness Optimization (AAO), a simple yet effective approach that automatically re-weights ambiguous content to reduce ambiguities by calculating semantic similarity from preference pairs. Through extensive experiments, we demonstrate that AAO consistently and significantly surpasses state-of-the-art approaches in performance, without markedly increasing response length, across multiple model scales and widely adopted benchmark datasets, including AlpacaEval 2, MT-Bench, and Arena-Hard. Specifically, AAO outperforms DPO by up to 8.9 points on AlpacaEval 2 and achieves an improvement of by up to 15.0 points on Arena-Hard.

Anthology ID:: 2025.emnlp-main.460
Volume:: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 9053–9063
Language:
URL:: https://aclanthology.org/2025.emnlp-main.460/
DOI:: 10.18653/v1/2025.emnlp-main.460
Bibkey:
Cite (ACL):: Jian Li, Shenglin Yin, Yujia Zhang, Alan Zhao, Xi Chen, Xiaohui Zhou, and Pengfei Xu. 2025. Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 9053–9063, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: Ambiguity Awareness Optimization: Towards Semantic Disambiguation for Direct Preference Optimization (Li et al., EMNLP 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.emnlp-main.460.pdf
Checklist:: 2025.emnlp-main.460.checklist.pdf

PDF Cite Search Checklist Fix data