Chen Huang
Author directoryOther people with similar names: Chen Huang
Unverified author pages with similar names: Chen Huang
2026
Towards Proactive Information Probing: Customer Service Chatbots Harvesting Value from Conversation
Chen Huang | Zitan Jiang | Zou Changyi | Wenqiang Lei | See-Kiong Ng
Findings of the Association for Computational Linguistics: ACL 2026
Chen Huang | Zitan Jiang | Zou Changyi | Wenqiang Lei | See-Kiong Ng
Findings of the Association for Computational Linguistics: ACL 2026
Customer service chatbots are increasingly expected to serve not merely as reactive support tools for users, but as strategic interfaces for harvesting high-value information and business intelligence. In response, we make three main contributions. 1) We introduce and define a novel task of Proactive Information Probing, which optimizes when to probe users for pre-specified target information while minimizing conversation turns and user friction. 2) We propose PROCHATIP, a proactive chatbot framework featuring a specialized conversation strategy module trained to master the delicate timing of probes. 3) Experiments demonstrate that PROCHATIP significantly outperforms baselines, exhibiting superior capability in both information probing and service quality. We believe that our work effectively redefines the commercial utility of chatbots, positioning them as scalable, cost-effective engines for proactive business intelligence. Our code is available at https://github.com/SCUNLP/PROCHATIP.
Beyond Prompt: Fine-grained Simulation of Cognitively Impaired Standardized Patients via Stochastic Steering
Weikang Zhang | Zimo Zhu | Zhichuan Yang | Chen Huang | Wenqiang Lei | See-Kiong Ng
Findings of the Association for Computational Linguistics: ACL 2026
Weikang Zhang | Zimo Zhu | Zhichuan Yang | Chen Huang | Wenqiang Lei | See-Kiong Ng
Findings of the Association for Computational Linguistics: ACL 2026
Simulating Standardized Patients with cognitive impairment offers a scalable and ethical solution for clinical training. However, existing methods rely on discrete prompt engineering and fail to capture the heterogeneity of deficits across varying domains and severity levels. To address this limitation, we propose StsPatient for the fine-grained simulation of cognitively impaired patients. We innovatively capture domain-specific features by extracting steering vectors from contrastive pairs of instructions and responses. Furthermore, we introduce a Stochastic Token Modulation (STM) mechanism to regulate the intervention probability. STM enables precise control over impairment severity while mitigating the instability of conventional vector methods. Comprehensive experiments demonstrate that StsPatient significantly outperforms baselines in both clinical authenticity and severity controllability. Our code will be open-sourced upon acceptance.
METRO: Towards Strategy Induction from Expert Dialogue Transcripts for Non-collaborative Dialogues
Haofu Yang | Jiaji Liu | Chen Huang | Faguo Wu | Wenqiang Lei | See-Kiong Ng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Haofu Yang | Jiaji Liu | Chen Huang | Faguo Wu | Wenqiang Lei | See-Kiong Ng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Developing non-collaborative dialogue agents traditionally requires the manual, unscalable codification of expert strategies. We propose METRO, a method that leverages large language models to autonomously induce both strategy actions and planning logic directly from raw transcripts. METRO formalizes expert knowledge into a Strategy Forest, a hierarchical structure that captures both short-term responses (nodes) and long-term strategic foresight (branches). Experimental results across two benchmarks show that METRO demonstrates promising performance, outperforming existing methods by an average of 9%-10%. Our further analysis not only reveals the success behind METRO (strategic behavioral diversity and foresight), but also demonstrates its robust cross-task transferability. This offers new insights into building non-collaborative agents in a cost-effective and scalable way. Our code is available at https://github.com/Humphrey-0125/METRO.
METER: Evaluating Multi-Level Contextual Causal Reasoning in Large Language Models
Pengfeng Li | Chen Huang | Chaoqun Hao | Hongyao Chen | Xiao-Yong Wei | Wenqiang Lei | See-Kiong Ng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Pengfeng Li | Chen Huang | Chaoqun Hao | Hongyao Chen | Xiao-Yong Wei | Wenqiang Lei | See-Kiong Ng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Contextual causal reasoning is a critical yet challenging capability for Large Language Models (LLMs). Existing benchmarks, however, often evaluate this skill in fragmented settings, failing to ensure context consistency or cover the full causal hierarchy. To address this, we pioneer METER to systematically benchmark LLMs across all three levels of the causal ladder under a unified context setting. Our extensive evaluation of various LLMs reveals a significant decline in proficiency as tasks ascend the causal hierarchy. To diagnose this degradation, we conduct a deep mechanistic analysis via both error pattern identification and internal information flow tracing. Our analysis reveals two primary failure modes: (1) LLMs are susceptible to distraction by causally irrelevant but factually correct information at lower level of causality; and (2) as tasks ascend the causal hierarchy, faithfulness to the provided context degrades, leading to a reduced performance. We belive our work advances our understanding of the mechanisms behind LLM contextual causal reasoning and establishes a critical foundation for future research. Our code and dataset are available at https://github.com/SCUNLP/METER.
E2EDev: Benchmarking Large Language Models in End-to-End Software Development Task
Jingyao Liu | Chen Huang | Zhizhao Guan | Wenqiang Lei | Yang Deng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Jingyao Liu | Chen Huang | Zhizhao Guan | Wenqiang Lei | Yang Deng
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD). However, existing E2ESD benchmarks are limited by coarse-grained requirement specifications and unreliable evaluation protocols, hindering a true understanding of current framework capabilities. To address these limitations, we present E2EDev, a novel benchmark grounded in the principles of Behavior-Driven Development (BDD) to assess whether the generated software meets user needs through mimicking real user interactions. E2EDev comprises (i) a fine-grained set of user requirements for each target software project (ii) multiple BDD test scenarios with corresponding Python step implementations for each requirement, and (iii) a fully automated testing pipeline built on the Behave framework. By evaluating various E2ESD frameworks and LLM backbones with E2EDev, our analysis reveals a persistent struggle to effectively solve these tasks, underscoring the critical need for more effective and cost-efficient E2ESD solutions. Our codebase and benchmark are available at https://github.com/SCUNLP/E2EDev.
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
Yifan Zhang | Chen Huang | Yueke Zhang | Jiahao Zhang | Toby Jia-Jun Li | Collin McMillan | Kevin Leach | Yu Huang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yifan Zhang | Chen Huang | Yueke Zhang | Jiahao Zhang | Toby Jia-Jun Li | Collin McMillan | Kevin Leach | Yu Huang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Code Language Models (CodeLLMs) traditionally learn attention based solely on statistical input-output token correlations (“machine attention”). In contrast, human developers rely on intuition, selectively fixating on semantically salient tokens during program comprehension. We present EyeMulator, a model-agnostic technique to align CodeLLM attention with human visual attention without architectural changes. By extracting scan paths from eye-tracking data, we derive token-level attention weights used to augment the loss function during fine-tuning. This induces the model to mimic human focus. Our evaluation across StarCoder, Llama-3.2, and DeepSeek-Coder shows that EyeMulator significantly outperforms baselines, achieving gains of over 30 CodeBLEU points in translation and up to 22 BERTScore points in summarization. Ablation studies confirm that these gains stem directly from replicating human attention dynamics. Artifacts are available at https://zenodo.org/records/17205682.
2025
Breaking the Stigma! Unobtrusively Probe Symptoms in Depression Disorder Diagnosis Dialogue
Jieming Cao | Chen Huang | Yanan Zhang | Ruibo Deng | Jincheng Zhang | Wenqiang Lei
Findings of the Association for Computational Linguistics: NAACL 2025
Jieming Cao | Chen Huang | Yanan Zhang | Ruibo Deng | Jincheng Zhang | Wenqiang Lei
Findings of the Association for Computational Linguistics: NAACL 2025
CANDY: Benchmarking LLMs’ Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
Ruiling Guo | Xinwei Yang | Chen Huang | Tong Zhang | Yong Hu
Findings of the Association for Computational Linguistics: EMNLP 2025
Ruiling Guo | Xinwei Yang | Chen Huang | Tong Zhang | Yong Hu
Findings of the Association for Computational Linguistics: EMNLP 2025
The effectiveness of large language models (LLMs) to fact-check misinformation remains uncertain, despite their growing use. To this end, we present CANDY, a benchmark designed to systematically evaluate the capabilities and limitations of LLMs in fact-checking Chinese misinformation. Specifically, we curate a carefully annotated dataset of ~20k instances. Our analysis shows that current LLMs exhibit limitations in generating accurate fact-checking conclusions, even when enhanced with chain-of-thought reasoning and few-shot prompting. To understand these limitations, we develop a taxonomy to categorize flawed LLM-generated explanations for their conclusions and identify factual fabrication as the most common failure mode. Although LLMs alone are unreliable for fact-checking, our findings indicate their considerable potential to augment human performance when deployed as assistive tools in scenarios. Our dataset and code can be accessed at https://github.com/SCUNLP/CANDY.
ASTRO: Automatic Strategy Optimization For Non-Cooperative Dialogues
Yikuan Hu | Chen Huang | Wenqiang Lei
Findings of the Association for Computational Linguistics: ACL 2025
Yikuan Hu | Chen Huang | Wenqiang Lei
Findings of the Association for Computational Linguistics: ACL 2025
Non-cooperative dialogues, such as negotiations and persuasion, present significant challenges for large language models (LLMs) due to the lack of inherent cooperation or shared goals. Current methods for optimizing dialogue strategies require substantial human effort for strategy optimization. To address these challenges, we propose ASTRO (Automated Strategy Optimization), a fully automated solution that leverages LLMs’ self-envolving capabilities. ASTRO dynamically generates customized strategy sets based on task goals and optimizes strategy planner using a self-play reinforcement learning paradigm. Our experimental results demonstrate ASTRO’s significant performance improvements over baseline models across various non-cooperative dialogue tasks, highlighting the potential for autonomously developing such agents without human intervention. Our code is available at https://github.com/SCUNLP/ASTRO.
Can Large Language Models Understand Internet Buzzwords Through User-Generated Content
Chen Huang | Junkai Luo | Xinzuo Wang | Wenqiang Lei | Jiancheng Lv
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chen Huang | Junkai Luo | Xinzuo Wang | Wenqiang Lei | Jiancheng Lv
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
The massive user-generated content (UGC) available in Chinese social media is giving rise to the possibility of studying internet buzzwords. In this paper, we study if large language models (LLMs) can generate accurate definitions for these buzzwords based on UGC as examples. Our work serves a threefold contribution. First, we introduce CHEER, the first dataset of Chinese internet buzzwords, each annotated with a definition and relevant UGC. Second, we propose a novel method, called RESS, to effectively steer the comprehending process of LLMs to produce more accurate buzzword definitions, mirroring the skills of human language learning. Third, with CHEER, we benchmark the strengths and weaknesses of various off-the-shelf definition generation methods and our RESS. Our benchmark demonstrates the effectiveness of RESS while revealing a crucial shared challenge: comprehending unseen buzzwords and leveraging sufficient, high-quality UGC to facilitate this comprehension. In this paper, we believe our work lays the groundwork for future advancements in LLM-based definition generation. Our dataset and code will be openly released.
ELABORATION: A Comprehensive Benchmark on Human-LLM Competitive Programming
Xinwei Yang | Zhaofeng Liu | Chen Huang | Jiashuai Zhang | Tong Zhang | Yifan Zhang | Wenqiang Lei
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xinwei Yang | Zhaofeng Liu | Chen Huang | Jiashuai Zhang | Tong Zhang | Yifan Zhang | Wenqiang Lei
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
While recent research increasingly emphasizes the value of human-LLM collaboration in competitive programming and proposes numerous empirical methods, a comprehensive understanding remains elusive due to the fragmented nature of existing studies and their use of diverse, application-specific human feedback. Thus, our work serves a three-fold purpose: First, we present the first taxonomy of human feedback consolidating the entire programming process, which promotes fine-grained evaluation. Second, we introduce ELABORATIONSET, a novel programming dataset specifically designed for human-LLM collaboration, meticulously annotated to enable large-scale simulated human feedback and facilitate cost-effective real human interaction studies. Third, we introduce ELABORATION, a novel benchmark to facilitate a thorough assessment of human-LLM competitive programming. With ELABORATION, we pinpoint strengthes and weaknesses of existing methods, thereby setting the foundation for furture improvement. Our dataset and code will be openly released.
How to Enable Effective Cooperation Between Humans and NLP Models: A Survey of Principles, Formalizations, and Beyond
Chen Huang | Yang Deng | Wenqiang Lei | Jiancheng Lv | Tat-Seng Chua | Jimmy Xiangji Huang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chen Huang | Yang Deng | Wenqiang Lei | Jiancheng Lv | Tat-Seng Chua | Jimmy Xiangji Huang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
With the advancement of large language models (LLMs), intelligent models have evolved from mere tools to autonomous agents with their own goals and strategies for cooperating with humans. This evolution has birthed a novel paradigm in NLP, i.e., human-model cooperation, that has yielded remarkable progress in numerous NLP tasks in recent years. In this paper, we take the first step to present a thorough review of human-model cooperation, exploring its principles, formalizations, and open challenges. In particular, we introduce a new taxonomy that provides a unified perspective to summarize existing approaches. Also, we discuss potential frontier areas and their corresponding challenges. We regard our work as an entry point, paving the way for more breakthrough research in this regard.
Cross-model Transferability among Large Language Models on the Platonic Representations of Concepts
Youcheng Huang | Chen Huang | Duanyu Feng | Wenqiang Lei | Jiancheng Lv
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Youcheng Huang | Chen Huang | Duanyu Feng | Wenqiang Lei | Jiancheng Lv
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Understanding the inner workings of Large Language Models (LLMs) is a critical research frontier. Prior research has shown that a single LLM’s concept representations can be captured as steering vectors (SVs), enabling the control of LLM behavior (e.g., towards generating harmful content). Our work takes a novel approach by exploring the intricate relationships between concept representations across different LLMs, drawing an intriguing parallel to Plato’s Allegory of the Cave. In particular, we introduce a linear transformation method to bridge these representations and present three key findings: 1) Concept representations across different LLMs can be effectively aligned using simple linear transformations, enabling efficient cross-model transfer and behavioral control via SVs. 2) This linear transformation generalizes across concepts, facilitating alignment and control of SVs representing different concepts across LLMs. 3) A weak-to-strong transferability exists between LLM concept representations, whereby SVs extracted from smaller LLMs can effectively control the behavior of larger LLMs. Our code is provided in the supplementary file and will be openly released.
Search
Fix author
Co-authors
- Wenqiang Lei 11
- See Kiong Ng 4
- Jiancheng Lv 3
- Yang Deng 2
- Xinwei Yang 2
- Tong Zhang 2
- Yifan Zhang 2
- Jieming Cao 1
- Zou Changyi 1
- Hongyao Chen 1
- Tat-Seng Chua 1
- Ruibo Deng 1
- Duanyu Feng 1
- Zhizhao Guan 1
- Ruiling Guo 1
- Chaoqun Hao 1
- Yikuan Hu 1
- Yong Hu 1
- Jimmy Xiangji Huang 1
- Youcheng Huang 1
- Yu Huang 1
- Zitan Jiang 1
- Kevin Leach 1
- Pengfeng Li 1
- Toby Jia-Jun Li 1
- Jiaji Liu 1
- Jingyao Liu 1
- Zhaofeng Liu 1
- Junkai Luo 1
- Collin McMillan 1
- Xinzuo Wang 1
- Xiao-Yong Wei 1
- Faguo Wu 1
- Haofu Yang 1
- Zhichuan Yang 1
- Jiahao Zhang 1
- Jiashuai Zhang 1
- Jincheng Zhang 1
- Weikang Zhang 1
- Yanan Zhang 1
- Yueke Zhang 1
- Zimo Zhu 1