Feng Chen
Author directoryOther people with similar names: Feng Chen, Feng Chen, Feng Chen
Unverified author pages with similar names: Feng Chen
2026
AgentOCR: Reimagining Agent History via Optical Self-Compression
Lang Feng | Fuchao Yang | Feng Chen | Xin Cheng | Haiyang Xu | Zhenglin Wan | Ming Yan | Bo An
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Lang Feng | Fuchao Yang | Feng Chen | Xin Cheng | Haiyang Xu | Zhenglin Wan | Ming Yan | Bo An
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction, but practical deployment is bottlenecked by rapidly growing textual histories that inflate token and memory costs. We introduce AgentOCR, a framework that exploits visual tokens’ superior information density by representing the accumulated observation-action history as a compact rendered image. To make multi-turn rollouts scalable, AgentOCR proposes segment optical caching. By decomposing history into hashable segments and maintaining a visual cache, this mechanism eliminates redundant re-rendering. Beyond fixed rendering, AgentOCR introduces agentic self-compression, where the agent actively emits a compression rate and is trained with compression-aware reward to adaptively balance task success and token efficiency. We conduct extensive experiments on challenging agentic benchmarks, ALFWorld and search-based QA. Remarkably, AgentOCR preserves over 95% of text-based agent performance while substantially reducing token consumption (>50%), yielding consistent token and memory efficiency. Further analysis validates a 20× rendering speedup from optical caching and effective self-compression balancing. Our code is available at https://github.com/langfengQ/AgentOCR.
2025
Lost in the Context: Insufficient and Distracted Attention to Contexts in Preference Modeling
Shihan Dou | Jiayi Chen | Chenhao Huang | Feng Chen | Wei Chengzhi | Huiyuan Zheng | Shichun Liu | Yan Liu | Chenxiao Liu | Chao Xin | Lin Yan | Zongzhang Zhang | Tao Gui | Qi Zhang | Xuanjing Huang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Shihan Dou | Jiayi Chen | Chenhao Huang | Feng Chen | Wei Chengzhi | Huiyuan Zheng | Shichun Liu | Yan Liu | Chenxiao Liu | Chao Xin | Lin Yan | Zongzhang Zhang | Tao Gui | Qi Zhang | Xuanjing Huang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
In Reinforcement Learning from Human Feedback (RLHF), the reward model (RM) evaluates the response quality based on the given context and assigns a reward. It plays a crucial role in aligning RLHF with human preferences. Although the current RM training paradigm concatenates the context and response while amplifying the reward difference between good and bad response pairs, we demonstrate that the RM faces two significant issues: i) it often allocates only a small proportion of attention to the context, and ii) it frequently ignores segments of the context that are relevant for evaluating the response quality. These issues undermine the RM’s effectiveness in modeling human preferences. To further address these challenges, we propose AttnRM, a novel optimization framework that enables the RM to concentrate on crucial segments of the context. Experimental results demonstrate that AttnRM significantly improves preference modeling by increasing attention to relevant information within the context. It also enhances the RM’s generalizability and achieves better performance in aligning with human preferences.