From Explicit to Implicit: A Theoretical Framework and Transfer Method for Preference Internalization in Language Models

Binrui Wang, Yongping Du, Yu Pei, Zikai Wang


Abstract
Transforming explicit preference signals into implicit and parameterized behaviors is pivotal for enabling prompt-free, human-aligned generation and improving the usability, efficiency and robustness of large language models. However, existing methods align model preference well but still rely on explicit user instructions to convey specific preferences, leading to cumbersome user experiences and undermining natural, frictionless interaction with the model. To fill the gap between explicit and implicit preference representation, this paper introduces a theoretical framework that establishes both necessary and sufficient conditions for effective preference recognition. Based on this framework, we propose a novel Two-Stage Progressive Preference Transfer (TSPPT) method, which decomposes preference internalization into two manageable stages: preference representation learning and preference internalization transfer. The proposed method fills the gap between explicit and implicit preferences while maintaining the model’s general capabilities. The experiments across multiple models (Qwen2.5, Qwen3, Llama-3.2, DeepSeek-R1-Distill) and datasets (UltraFeedback, HelpSteer) demonstrate superior performance. The proposed method achieves 79.2% win rate on UltraFeedback (vs. 59.2–67.6% for baselines), substantial improvements on MT-Bench (7.86 vs. 7.34 for best baseline), and significant reductions in implicit social bias (0.165 vs. 0.185–0.325 for baselines). Notably, the method maintains comparable performance between implicit and explicit settings, confirming successful preference internalization.1
Anthology ID:
2026.tacl-1.48
Volume:
Transactions of the Association for Computational Linguistics, Volume 14
Month:
Year:
2026
Address:
Cambridge, MA
Venue:
TACL
SIG:
Publisher:
MIT Press
Note:
Pages:
1074–1095
Language:
URL:
https://aclanthology.org/2026.tacl-1.48/
DOI:
10.1162/tacl.a.697
Bibkey:
Cite (ACL):
Binrui Wang, Yongping Du, Yu Pei, and Zikai Wang. 2026. From Explicit to Implicit: A Theoretical Framework and Transfer Method for Preference Internalization in Language Models. Transactions of the Association for Computational Linguistics, 14:1074–1095.
Cite (Informal):
From Explicit to Implicit: A Theoretical Framework and Transfer Method for Preference Internalization in Language Models (Wang et al., TACL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.tacl-1.48.pdf