Fei Huang
Other people with similar names: Fei Huang, Fei Huang, Fei Huang
Unverified author pages with similar names: Fei Huang
2026
ToolRM: Towards Agentic Tool-Use Reward Modeling
Renhao Li | Jianhong Tu | Yang Su | Yantao Liu | Fei Huang | Hamid Alinejad-Rokny | Derek F. Wong | Junyang Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Renhao Li | Jianhong Tu | Yang Su | Yantao Liu | Fei Huang | Hamid Alinejad-Rokny | Derek F. Wong | Junyang Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Reward models (RMs) play a critical role in aligning large language models (LLMs) with human preferences. Yet in the domain of tool learning, the lack of RMs specifically designed for function-calling tasks has limited progress toward more capable agentic AI. We introduce ToolRM, a family of lightweight reward models tailored for general tool-use scenarios. To build these models, we propose a novel pipeline that constructs high-quality pairwise preference data using rule-based scoring and multidimensional sampling. This yields ToolPref-Pairwise-30K, a diverse, balanced, and challenging preference dataset that supports both generative and discriminative reward modeling. We also introduce TRBenchBFCL, a benchmark built on the agent evaluation suite BFCL to evaluate RMs on tool calling tasks. Trained on our constructed data, models from the Qwen3-4B/8B series achieve up to 17.94% higher accuracy, substantially outperforming frontier LLMs and RMs in pairwise reward judgments. Beyond training objectives, generative ToolRM generalizes to broader critique tasks, including Best-of-N sampling and self-correction. Experiments on ACEBench highlight its effectiveness and efficiency, enabling inference-time scaling while reducing output token usage by over 66%. Its support for downstream RL training further validates its practical utility. We release data to facilitate future research.
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
Binghai Wang | Yantao Liu | Yuxuan Liu | Tianyi Tang | Shenzhi Wang | Chang Gao | Chujie Zheng | Yichang Zhang | Le Yu | Shixuan Liu | Tao Gui | Qi Zhang | Xuanjing Huang | Bowen Yu | Fei Huang | Junyang Lin
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Binghai Wang | Yantao Liu | Yuxuan Liu | Tianyi Tang | Shenzhi Wang | Chang Gao | Chujie Zheng | Yichang Zhang | Le Yu | Shixuan Liu | Tao Gui | Qi Zhang | Xuanjing Huang | Bowen Yu | Fei Huang | Junyang Lin
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Generative Reward Models (GenRMs) and LLM-as-a-Judge exhibit deceptive alignment by producing correct judgments for incorrect reasons, as they are trained and evaluated to prioritize Outcome Accuracy, which undermines their ability to generalize during RLHF. We introduce Rationale Consistency, a fine-grained metric that quantifies the alignment between the model’s reasoning process and human judgment. Our evaluation of frontier models reveals that rationale consistency effectively discriminates among state-of-the-art models and detects deceptive alignment, while outcome accuracy falls short in both respects. To mitigate this gap, we introduce a hybrid signal that combines rationale consistency with outcome accuracy for GenRM training. Our training method achieves state-of-the-art performance on RM-Bench (87.1%) and JudgeBench (82%), surpassing outcome-only baselines by an average of 5%. Using RM during RLHF, our method effectively improves performance as demonstrated on Arena Hard v2, notably yielding a 7% improvement in creative writing tasks. Further analysis confirms that our method escapes the deceptive alignment trap, effectively reversing the decline in rationale consistency observed in outcome-only training.
2025
RMTBench: Benchmarking LLMs Through Multi-Turn User-Centric Role-Playing
Hao Xiang | Tianyi Tang | Yang Su | Bowen Yu | An Yang | Fei Huang | Yichang Zhang | Yaojie Lu | Hongyu Lin | Xianpei Han | Jingren Zhou | Junyang Lin | Le Sun
Findings of the Association for Computational Linguistics: EMNLP 2025
Hao Xiang | Tianyi Tang | Yang Su | Bowen Yu | An Yang | Fei Huang | Yichang Zhang | Yaojie Lu | Hongyu Lin | Xianpei Han | Jingren Zhou | Junyang Lin | Le Sun
Findings of the Association for Computational Linguistics: EMNLP 2025
Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly adopt a character-centric approach, simplify user-character interactions to isolated Q&A tasks, and fail to reflect real-world applications. To address this limitation, we introduce RMTBench, a comprehensive user-centric bilingual role-playing benchmark featuring 80 diverse characters and over 8,000 dialogue rounds. RMTBench includes custom characters with detailed backgrounds and abstract characters defined by simple traits, enabling evaluation across various user scenarios. Our benchmark constructs dialogues based on explicit user motivations rather than character descriptions, ensuring alignment with practical user applications. Furthermore, we construct an authentic multi-turn dialogue simulation mechanism. With carefully selected evaluation dimensions and LLM-based scoring, this mechanism captures the complex intention of conversations between the user and the character. By shifting focus from character background to user intention fulfillment, RMTBench bridges the gap between academic evaluation and practical deployment requirements, offering a more effective framework for assessing role-playing capabilities in LLMs. All code and datasets will be released soon.
LongWeave: A Long-Form Generation Benchmark Bridging Real-World Relevance and Verifiability
Zikai Xiao | Fei Huang | Jianhong Tu | Jianhui Wei | Wen Ma | Yuxuan Zhou | Jian Wu | Bowen Yu | Zuozhu Liu | Junyang Lin
Findings of the Association for Computational Linguistics: EMNLP 2025
Zikai Xiao | Fei Huang | Jianhong Tu | Jianhui Wei | Wen Ma | Yuxuan Zhou | Jian Wu | Bowen Yu | Zuozhu Liu | Junyang Lin
Findings of the Association for Computational Linguistics: EMNLP 2025
Generating long, informative, and factual outputs remains a major challenge for Large Language Models (LLMs). Existing benchmarks for long-form generation typically assess real-world queries with hard-to-verify metrics or use synthetic setups that ease evaluation but overlook real-world intricacies. In this paper, we introduce LongWeave, which balance real-world and verifiable assessment with Target-Anchored Evaluation (TAE). TAE constructs tasks by first defining verifiable targets within real-world scenarios, then systematically generating corresponding queries, textual materials, and anchors based on these targets. This ensures that tasks are both realistic and objectively assessable, enabling rigorous assessment of model capabilities in meeting complex real-world constraints. LongWeave supports customizable input/output lengths (up to 64K/8K tokens) across seven distinct tasks. Evaluation on 23 LLMs show that even state-of-the-art models encounter significant challenges in long-form generation as real-world complexity and output length increase. Dataset will be publicly available.
NOVA-63: Native Omni-lingual Versatile Assessments of 63 Disciplines
Jinyang Zhang | Kexin Yang | Yu Wan | Muyang Ye | Yue Fang | Baosong Yang | Fei Huang | Junyang Lin | Dayiheng Liu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jinyang Zhang | Kexin Yang | Yu Wan | Muyang Ye | Yue Fang | Baosong Yang | Fei Huang | Junyang Lin | Dayiheng Liu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
The multilingual capabilities of large language models (LLMs) have attracted considerable attention over the past decade. Assessing the accuracy with which LLMs provide answers in multilingual contexts is essential for determining their level of multilingual proficiency. Nevertheless, existing multilingual benchmarks generally reveal severe drawbacks, such as overly translated content (translationese), the absence of difficulty control, constrained diversity, and disciplinary imbalance, making the benchmarking process unreliable and showing low convincingness. To alleviate those shortcomings, we introduce NOVA-63 (Native Omni-lingual Versatile Assessments of 63 Disciplines), a comprehensive, difficult multilingual benchmark featuring 93,536 questions sourced from native speakers across 14 languages and 63 academic disciplines. Leveraging a robust pipeline that integrates LLM-assisted formatting, expert quality verification, and multi-level difficulty screening, NOVA-63 is balanced on disciplines with consistent difficulty standards while maintaining authentic linguistic elements. Extensive experimentation with current LLMs has shown significant insights into cross-lingual consistency among language families, and exposed notable disparities in models’ capabilities across various disciplines. This work provides valuable benchmarking data for the future development of multilingual models. Furthermore, our findings underscore the importance of moving beyond overall scores and instead conducting fine-grained analyses of model performance.
P-MMEval: A Parallel Multilingual Multitask Benchmark for Consistent Evaluation of LLMs
Yidan Zhang | Yu Wan | Boyi Deng | Baosong Yang | Hao-Ran Wei | Fei Huang | Bowen Yu | Dayiheng Liu | Junyang Lin | Fei Huang | Jingren Zhou
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Yidan Zhang | Yu Wan | Boyi Deng | Baosong Yang | Hao-Ran Wei | Fei Huang | Bowen Yu | Dayiheng Liu | Junyang Lin | Fei Huang | Jingren Zhou
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Recent advancements in large language models (LLMs) showcase varied multilingual capabilities across tasks like translation, code generation, and reasoning. Previous assessments often limited their scope to fundamental natural language processing (NLP) or isolated capability-specific tasks. To alleviate this drawback, we aim to present a comprehensive multilingual multitask benchmark. First, we introduce P-MMEval, a large-scale benchmark covering fundamental and capability-specialized datasets. Furthermore, P-MMEval delivers consistent language coverage across various datasets and provides parallel samples. Finally, we conduct extensive experiments on representative multilingual model series to compare performances across models and tasks, explore the relationship between multilingual performances and factors such as tasks, model sizes, languages, and prompts, and examine the effectiveness of knowledge transfer from English to other languages. The resulting insights are intended to offer valuable guidance for future research.
Search
Fix author
Co-authors
- Junyang Lin 6
- Bowen Yu 4
- Dayiheng Liu 2
- Yantao Liu 2
- Yang Su 2
- Tianyi Tang 2
- Jianhong Tu 2
- Yu Wan 2
- Baosong Yang 2
- Yichang Zhang 2
- Jingren Zhou 2
- Hamid Alinejad-Rokny 1
- Boyi Deng 1
- Yue Fang 1
- Chang Gao 1
- Tao Gui 1
- Xianpei Han 1
- Fei Huang 1
- Xuan-Jing Huang (黄萱菁) 1
- Renhao Li 1
- Hongyu Lin 1
- Shixuan Liu (刘世萱) 1
- Yuxuan Liu 1
- Zuozhu Liu 1
- Yaojie Lu 1
- Wen Ma 1
- Le Sun 1
- Binghai Wang 1
- Shenzhi Wang 1
- Hao-Ran Wei 1
- Jianhui Wei 1
- Derek F. Wong (黄辉) 1
- Jian Wu 1
- Hao Xiang 1
- Zikai Xiao 1
- An Yang 1
- Kexin Yang 1
- Min Yang 1
- Muyang Ye 1
- Le Yu 1
- Jinyang Zhang 1
- Qi Zhang 1
- Yidan Zhang 1
- Chujie Zheng 1
- Yuxuan Zhou 1