Min Yang
Other people with similar names: Min Yang, Min Yang, Min Yang
Unverified author pages with similar names: Min Yang
2026
Beyond Quantity: Trajectory Diversity Scaling for Code Agents
Guhong Chen | Chenghao Sun | Cheng Fu | Qiyao Wang | Zhihong Huang | ChaoPeng Wei | Guangxu Chen | Feiteng Fang | Ahmadreza Argha | Bing Zhao | Xander Xu | Qi Han | Hamid Alinejad-Rokny | Qiang Qu | Binhua Li | Shiwen Ni | Min Yang | HU Wei | Yongbin Li
Findings of the Association for Computational Linguistics: ACL 2026
Guhong Chen | Chenghao Sun | Cheng Fu | Qiyao Wang | Zhihong Huang | ChaoPeng Wei | Guangxu Chen | Feiteng Fang | Ahmadreza Argha | Bing Zhao | Xander Xu | Qi Han | Hamid Alinejad-Rokny | Qiang Qu | Binhua Li | Shiwen Ni | Min Yang | HU Wei | Yongbin Li
Findings of the Association for Computational Linguistics: ACL 2026
As code large language models (LLMs) evolve into tool-interactive agents via the Model Context Protocol (MCP), their generalization is increasingly limited by low-quality synthetic data and the diminishing returns of quantity scaling; moreover, quantity-centric scaling exhibits an early bottleneck that underutilizes trajectory data. We propose TDScaling, a Trajectory Diversity Scaling-based data synthesis framework for code agents that scales performance through diversity rather than raw volume. Moreover, TDScaling is more data-efficient: under a fixed training budget, increasing trajectory diversity yields larger gains than adding more trajectories, improving the performance-cost trade-off for agent training. TDScaling integrates four innovations: (1) a Business Cluster mechanism that captures real-service logical dependencies; (2) a Blueprint-driven multi-agent paradigm that enforces trajectory coherence; (3) an adaptive evolution mechanism that steers synthesis toward long-tail scenarios using Domain Entropy, Reasoning Mode Entropy, and Cumulative Action Complexity to prevent mode collapse; and (4) a sandboxed code tool that mitigates catastrophic forgetting of intrinsic coding capabilities. Experiments on general tool-use benchmarks (BFCL, š2-Bench) and code agent tasks (RebenchT, CodeCI, BIRD) demonstrate a win-win outcome: TDScaling improves both tool-use generalization and inherent coding proficiency. Crucially, we show that trajectory diversity scaling attains a substantially higher performance ceiling than quantity scaling, establishing a resource-efficient paradigm for training robust code agents under data bottlenecks.
ToolRM: Towards Agentic Tool-Use Reward Modeling
Renhao Li | Jianhong Tu | Yang Su | Yantao Liu | Fei Huang | Hamid Alinejad-Rokny | Derek F. Wong | Junyang Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Renhao Li | Jianhong Tu | Yang Su | Yantao Liu | Fei Huang | Hamid Alinejad-Rokny | Derek F. Wong | Junyang Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Reward models (RMs) play a critical role in aligning large language models (LLMs) with human preferences. Yet in the domain of tool learning, the lack of RMs specifically designed for function-calling tasks has limited progress toward more capable agentic AI. We introduce ToolRM, a family of lightweight reward models tailored for general tool-use scenarios. To build these models, we propose a novel pipeline that constructs high-quality pairwise preference data using rule-based scoring and multidimensional sampling. This yields ToolPref-Pairwise-30K, a diverse, balanced, and challenging preference dataset that supports both generative and discriminative reward modeling. We also introduce TRBenchBFCL, a benchmark built on the agent evaluation suite BFCL to evaluate RMs on tool calling tasks. Trained on our constructed data, models from the Qwen3-4B/8B series achieve up to 17.94% higher accuracy, substantially outperforming frontier LLMs and RMs in pairwise reward judgments. Beyond training objectives, generative ToolRM generalizes to broader critique tasks, including Best-of-N sampling and self-correction. Experiments on ACEBench highlight its effectiveness and efficiency, enabling inference-time scaling while reducing output token usage by over 66%. Its support for downstream RL training further validates its practical utility. We release data to facilitate future research.
SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models
Shuaimin Li | Liyang Fan | Zeyang li | Zhuoyue Wan | Yufang Lin | Shiwen Ni | Feiteng Fang | Hamid Alinejad-Rokny | Yuanfeng Song | Kun Jing | Chen Jason Zhang | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Shuaimin Li | Liyang Fan | Zeyang li | Zhuoyue Wan | Yufang Lin | Shiwen Ni | Feiteng Fang | Hamid Alinejad-Rokny | Yuanfeng Song | Kun Jing | Chen Jason Zhang | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by exposure to benchmark data during pre-training. Existing approaches either assume access to proprietary training corpora, rely on brittle heuristics such as timestamp filtering, or use external reference sets with manually tuned, non-generalizable thresholds. To address these limitations, we introduce SrDetection, a unified self-referential leakage detection framework for both gray-box (access to model logits) and black-box (access to model outputs) settings. SrDetection generates semantically equivalent variants of a benchmark sample and detects leakage by contrasting the modelās behavior on the original versus its variants, flagging cases where the original is disproportionately easier for the model. We further design a controlled leakage detection testbed and evaluate SrDetection in this environment. Across different models and training stages, SrDetection improves average F1 by 21.52 points in the gray-box setting and 14.46 points in the black-box setting over strong baselines, demonstrating robust, threshold-independent leakage detection. Finally, a gray-box study of 15 widely used Code LLMs on four popular benchmarks reveals benchmark-specific leakage patterns beyond prior overlap-based analyses[Source code and data are available at https://github.com/SMinL/SrDetectionCode].
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
Siyi Li | Jiajun Shi | Shiwen Ni | Ge Zhang | Shuaimin Li | Shijian Wang | Zhoufutu Wen | Yizhi LI | Hamid Alinejad-Rokny | Jiaheng Liu | Min Yang | Wenhao Huang
Findings of the Association for Computational Linguistics: ACL 2026
Siyi Li | Jiajun Shi | Shiwen Ni | Ge Zhang | Shuaimin Li | Shijian Wang | Zhoufutu Wen | Yizhi LI | Hamid Alinejad-Rokny | Jiaheng Liu | Min Yang | Wenhao Huang
Findings of the Association for Computational Linguistics: ACL 2026
Large Reasoning Models (LRMs) have demonstrated strong performance by producing extended Chain-of-Thought (CoT) traces before answering. However, this paradigm often induces over-reasoning: redundant calculations and circular self-verification that increase computational cost without improving outcomes. Existing evaluations largely emphasize final accuracy or coarse token counts, and lack automated tools to separate essential logic from structural redundancy. We introduce CoTJudger, a graph-driven framework that quantifies reasoning efficiency by converting free-form CoTs into directed dependency graphs and extracting the Shortest Effective Path (SEP) needed to reach a correct solution. This yields an interpretable efficiency signal ā how much of a CoT is necessary versus structurally redundant ā that is comparable across models and tasks. Evaluating 21 LRMs, CoTJudger reveals pervasive redundancy and surfaces recurring failure modes, including verification obsession and compensatory redundancy. These results provide a practical metric for disentangling reasoning ability from computational waste, enabling more targeted evaluation and diagnosis of LRM efficiency.
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning
Ruijun Huang | Zhiqiao Kang | Yuxuan Zhu | Junxiong Li | Jiahao Zhao | Minghuan Tan | Feng Jiang | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Ruijun Huang | Zhiqiao Kang | Yuxuan Zhu | Junxiong Li | Jiahao Zhao | Minghuan Tan | Feng Jiang | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
The accurate extraction of scientific measurements from literature is a critical yet challenging task in AI4Science, enabling large-scale analysis and integration of quantitative research findings. However, Large Language Models (LLMs) frequently exhibit severe hallucinations, which significantly undermine the reliability of automated scientific document understanding systems. To address this problem, we propose MeasHalu, a novel framework for mitigating scientific measurement hallucinations through enhanced reasoning and targeted optimization. We first present a fine-grained taxonomy of measurement-specific hallucinations, categorizing errors across quantities, units, modifiers, and relations. Our approach incorporates a two-stage reasoning-aware fine-tuning strategy using augmented scientific data and process-based supervision. Furthermore, we introduce a progressive reward curriculum designed to penalize specific hallucination types, significantly improving extraction faithfulness. Experimental results demonstrate that MeasHalu substantially reduces hallucination rates and improves overall accuracy on the MeasEval benchmark. This work provides a targeted solution to a key bottleneck in automated scientific knowledge extraction, facilitating more trustworthy and scalable machine-assisted scientific literature analysis.
Towards IP Intelligence: Benchmarking Large Language Models on Intellectual Property Knowledge and Practice
Qiyao Wang | Guhong Chen | Hongbo Wang | Huaren Liu | Minghui Zhu | Zhifei Qin | Li Linwei | Yilin Yue | Shiqiang Wang | Jiayan Li | Wu Yihang | Ziqiang Liu | Longze Chen | Run Luo | Liyang Fan | Jiaming Li | Lei Zhang | Kan Xu | Hamid Alinejad-Rokny | Chengming Li | Shiwen Ni | Yuan Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Qiyao Wang | Guhong Chen | Hongbo Wang | Huaren Liu | Minghui Zhu | Zhifei Qin | Li Linwei | Yilin Yue | Shiqiang Wang | Jiayan Li | Wu Yihang | Ziqiang Liu | Longze Chen | Run Luo | Liyang Fan | Jiaming Li | Lei Zhang | Kan Xu | Hamid Alinejad-Rokny | Chengming Li | Shiwen Ni | Yuan Lin | Min Yang
Findings of the Association for Computational Linguistics: ACL 2026
Intellectual Property (IP) is a highly specialized domain that integrates technical and legal knowledge, making it inherently complex and knowledge-intensive. Recent advancements in LLMs have demonstrated their potential to handle IP tasks, enabling more efficient analysis, understanding, and generation of IP-related content. However, existing datasets and benchmarks focus narrowly on patents or cover limited aspects of the IP field, lacking alignment with real-world scenarios. To bridge this gap, we introduce IPBench, the first comprehensive IP task taxonomy and a large-scale bilingual benchmark encompassing 8 IP mechanisms and 20 distinct tasks, designed to evaluate LLMs in real-world IP practice. We benchmark 19 main LLMs, ranging from general purpose to domain-specific, including chat-oriented and reasoning-focused models, under zero-shot, few-shot, and chain-of-thought settings. Our results show that even the top-performing model, DeepSeek-V3, achieves only 75.8% accuracy, indicating significant room for improvement. Notably, open-source IP and law-oriented models lag behind closed-source general-purpose models. To foster future research, we publicly release IPBench, and will expand it with additional tasks to better reflect real-world complexities and support model advancements in the IP domain. We provide the data, code in the supplementary materials.
Act-Adaptive Margin: Dynamically Calibrating Reward Models for Subjective Ambiguity
Feiteng Fang | Dingwei Chen | Xiang Huang | Ting-En Lin | Yuchuan Wu | Xiong Liu | Jing Ye | Ziqiang Liu | Haonan Zhang | Liang Zhu | Hamid Alinejad-Rokny | Min Yang | Yongbin Li
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Feiteng Fang | Dingwei Chen | Xiang Huang | Ting-En Lin | Yuchuan Wu | Xiong Liu | Jing Ye | Ziqiang Liu | Haonan Zhang | Liang Zhu | Hamid Alinejad-Rokny | Min Yang | Yongbin Li
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Currently, most reinforcement learning tasks focus on domains like mathematics and programming, where verification is relatively straightforward. However, in subjective tasks such as role-playing, alignment techniques struggle to make progress, primarily because subjective reward modeling using the Bradley-Terry model faces significant challenges when dealing with ambiguous preferences. To improve reward modeling in subjective tasks, this paper proposes AAM (Act-Adaptive Margin), which enhances reward modeling by dynamically calibrating preference margins using the modelās internal parameter knowledge. We design two versions of AAM that efficiently generate contextually-appropriate preference gaps without additional human annotation. This approach fundamentally improves how reward models handle subjective rewards by better integrating generative understanding with preference scoring. To validate AAMās effectiveness in subjective reward modeling, we conduct evaluations on RewardBench, JudgeBench, and challenging role-playing tasks. Results show that AAM significantly improves subjective reward modeling performance, enhancing Bradley-Terry reward models by 2.95% in general tasks and 4.85% in subjective role-playing tasks. Furthermore, reward models trained with AAM can help downstream alignment tasks achieve better results. Our test results show that applying rewards generated by AAM-Augmented RM to preference learning techniques (e.g., GRPO) achieves state-of-the-art results on CharacterEval and Charm. The code and dataset will be released upon acceptance.
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
Qiancheng Xu | Yongqi Li | Fan Liu | Hongru Wang | Min Yang | Wenjie Li
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Qiancheng Xu | Yongqi Li | Fan Liu | Hongru Wang | Min Yang | Wenjie Li
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Modelsā (LLMs) capabilities with external tools during reasoning. Existing TIR methods typically rely on external tool documentation during reasoning. However, this leads to tool mastery difficulty, tool size constraints, and inference inefficiency. To mitigate these issues, we explore Tool-Internalized Reasoning (TInR), aiming at facilitating reasoning with tool knowledge internalized into LLMs. Achieving this goal presents notable requirements, including tool internalization and tool-reasoning coordination. To address them, we propose TInR-U, a tool-internalized reasoning framework for unified reasoning and tool usage. TInR-U is trained through a three-phase pipeline: 1) tool internalization with a bidirectional knowledge alignment strategy; 2) supervised fine-tuning warm-up using high-quality reasoning annotations, and 3) reinforcement learning with TInR-specific rewards. We comprehensively evaluate our method across in-domain and out-of-domain settings. Experiment results show that TInR-U achieves superior performance in both settings, highlighting its effectiveness and efficiency. The codes are attached in the supplementary file for review.
Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers
Xin Chen | Feng Jiang | Yiqian Zhang | Hardy Chen | Shuo Yan | Wenya Xie | Min Yang | Shujian Huang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xin Chen | Feng Jiang | Yiqian Zhang | Hardy Chen | Shuo Yan | Wenya Xie | Min Yang | Shujian Huang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Reasoning-oriented Large Language Models (LLMs) have achieved remarkable progress with Chain-of-Thought (CoT) prompting, yet they remain fundamentally limited by a blind self-thinking paradigm: performing extensive internal reasoning even when critical information is missing or ambiguous. We propose Proactive Interactive Reasoning (PIR), a new reasoning paradigm that transforms LLMs from passive solvers into proactive inquirers that interleave reasoning with clarification. Unlike existing search- or tool-based frameworks that primarily address knowledge uncertainty by querying external environments, PIR targets premise- and intent-level uncertainty through direct interaction with the user. PIR is implemented via two core components: (1) an uncertainty-aware supervised fine-tuning procedure that equips models with interactive reasoning capability, and (2) a user-simulator-based policy optimization framework driven by a composite reward that aligns model behavior with user intent. Extensive experiments on mathematical reasoning, code generation, and document editing demonstrate that PIR consistently outperforms strong baselines, achieving up to 32.70% higher accuracy, 22.90% higher pass rate, and 41.36 BLEU improvement, while reducing nearly half of the reasoning computation and unnecessary interaction turns. Further reliability evaluations on factual knowledge, question answering, and missing-premise scenarios confirm the strong generalization and robustness of PIR.
2025
COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning
Yuelin Bai | Xeron Du | Yiming Liang | Leo Jin | Junting Zhou | Ziqiang Liu | Feiteng Fang | Mingshan Chang | Tianyu Zheng | Xincheng Zhang | Nuo Ma | Zekun Moore Wang | Ruibin Yuan | Haihong Wu | Hongquan Lin | Wenhao Huang | Jiajun Zhang | Chenghua Lin | Jie Fu | Min Yang | Shiwen Ni | Ge Zhang
Findings of the Association for Computational Linguistics: NAACL 2025
Yuelin Bai | Xeron Du | Yiming Liang | Leo Jin | Junting Zhou | Ziqiang Liu | Feiteng Fang | Mingshan Chang | Tianyu Zheng | Xincheng Zhang | Nuo Ma | Zekun Moore Wang | Ruibin Yuan | Haihong Wu | Hongquan Lin | Wenhao Huang | Jiajun Zhang | Chenghua Lin | Jie Fu | Min Yang | Shiwen Ni | Ge Zhang
Findings of the Association for Computational Linguistics: NAACL 2025
Remarkable progress on large language models (LLMs), particularly in English, has facilitated impressive capabilities in following human instructions. However, there remains a noticeable gap in instruction fine-tuning for Chinese, where the complex linguistic features pose significant challenges. Existing datasets, generally distilled from English-centric LLMs, are not well-aligned with Chinese usersā interaction patterns. To bridge this gap, we introduce COIG-CQIA, a new Chinese instruction tuning dataset derived from various real-world data resources and undergoing comprehensive human verification. We conduct extensive experiments on COIG-CQIA, and compare them with strong baseline models and datasets. The experimental results show that models trained on COIG-CQIA achieve highly competitive performance in diverse benchmarks. Additionally, our findings offer several insights for designing effective Chinese instruction-tuning datasets and data mixing strategies. Our dataset are available at https://huggingface.co/datasets/m-a-p/COIG-CQIA.
Unsupervised Speech-text word-level alignment with Dynamic Programming
Tianshu Yu | Zihan Gong | Minghuan Tan | Guhong Chen | Min Yang
Findings of the Association for Computational Linguistics: NAACL 2025
Tianshu Yu | Zihan Gong | Minghuan Tan | Guhong Chen | Min Yang
Findings of the Association for Computational Linguistics: NAACL 2025
HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluation
Shijie Zhang | Renhao Li | Songsheng Wang | Philipp Koehn | Min Yang | Derek F. Wong
Findings of the Association for Computational Linguistics: EMNLP 2025
Shijie Zhang | Renhao Li | Songsheng Wang | Philipp Koehn | Min Yang | Derek F. Wong
Findings of the Association for Computational Linguistics: EMNLP 2025
The advancement of Large Language Models (LLMs) enables flexible and interpretable automatic evaluations. In the field of machine translation evaluation, utilizing LLMs with translation error annotations based on Multidimensional Quality Metrics (MQM) yields more human-aligned judgments. However, current LLM-based evaluation methods still face challenges in accurately identifying error spans and assessing their severity. In this paper, we propose HiMATE, a Hierarchical Multi-Agent Framework for Machine Translation Evaluation. We argue that existing approaches inadequately exploit the fine-grained structural and semantic information within the MQM hierarchy. To address this, we develop a hierarchical multi-agent system grounded in the MQM error typology, enabling granular evaluation of subtype errors. Two key strategies are incorporated to further mitigate systemic hallucinations within the framework: the utilization of the modelās self-reflective capability and the facilitation of agent discussion involving asymmetric information. Empirically, HiMATE outperforms competitive baselines across different datasets in conducting human-aligned evaluations. Further analyses underscore its significant advantage in error span detection and severity assessment, achieving an average F1-score improvement of 89% over the best-performing baseline. We make our code and data publicly available at https://github.com/nlp2ct-shijie/HiMATE.
Forget for Get: A Lightweight Two-phase Gradient Method for Knowledge Editing in Large Language Models
Yanhong Li | Min Yang | Xiping Hu | Chengming Li
Findings of the Association for Computational Linguistics: EMNLP 2025
Yanhong Li | Min Yang | Xiping Hu | Chengming Li
Findings of the Association for Computational Linguistics: EMNLP 2025
Recent studies have highlighted the remarkable knowledge retention capabilities of Large Language Models (LLMs) like GPT-4, while simultaneously revealing critical limitations in maintaining knowledge currency and accuracy. Existing knowledge editing methodologies, designed to update specific factual information without compromising general model performance, often encounter two fundamental challenges: parameter conflict during knowledge overwriting and excessive computational overhead. In this paper, we introduce ForGet (Forget for Get), a novel approach grounded in the principle of āforgetting before learningā. By pinpointing the location within the LLM that corresponds to the target knowledge, we first erase the outdated knowledge and then insert the new knowledge at this precise spot. ForGet is the first work to leverage a two-phase gradient-based process for knowledge editing, offering a lightweight solution that also delivers superior results. Experimental findings show that our method achieves more effective knowledge editing at a lower cost compared to previous techniques across various base models.
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework
Yupeng Qi | Ziyu Lyu | Min Yang | Yanlin Wang | Lu Bai | Lixin Cui
Findings of the Association for Computational Linguistics: EMNLP 2025
Yupeng Qi | Ziyu Lyu | Min Yang | Yanlin Wang | Lu Bai | Lixin Cui
Findings of the Association for Computational Linguistics: EMNLP 2025
As large language models (LLMs) are increasingly applied across various domains, enhancing safety while maintaining the helpfulness of LLMs has become a critical challenge. Recent studies solve this problem through safety-constrained online preference optimization or safety-constrained offline preference optimization. However, the safety-constrained online methods often suffer from excessive safety, which might reduce helpfulness, while the safety-constrained offline methods perform poorly in adaptively balancing safety and helpfulness. To address these limitations, we propose MidPO, a Mixture of Experts (MoE) framework for safety-helpfulness dual Preference Optimization. Firstly, MidPO devises single-preference enhanced direct preference optimization approach to transform the base model into two independent experts, termed safety and helpfulness experts, and fine-tunes the two independent experts for optimal safety or helpfulness performance. Secondly, to achieve an effective balance between safety and helpfulness, MidPO incorporates the two experts into the MoE framework and designs a dynamic routing mechanism to allocate contributions from each expert adaptively. We conduct quantitative and qualitative experiments on three popular datasets to demonstrate the proposed MidPO significantly outperforms state-of-the-art approaches in both safety and helpfulness. Code is available at https://github.com/OutdoorManofML/MidPO.
LIME: Less Is More for MLLM Evaluation
King Zhu | Qianbo Zang | Shian Jia | Siwei Wu | Feiteng Fang | Yizhi Li | Shuyue Guo | Tianyu Zheng | Jiawei Guo | Bo Li | Haoning Wu | Xingwei Qu | Jian Yang | Ruibo Liu | Xiang Yue | Jiaheng Liu | Chenghua Lin | Hamid Alinejad-Rokny | Min Yang | Shiwen Ni | Wenhao Huang | Ge Zhang
Findings of the Association for Computational Linguistics: ACL 2025
King Zhu | Qianbo Zang | Shian Jia | Siwei Wu | Feiteng Fang | Yizhi Li | Shuyue Guo | Tianyu Zheng | Jiawei Guo | Bo Li | Haoning Wu | Xingwei Qu | Jian Yang | Ruibo Liu | Xiang Yue | Jiaheng Liu | Chenghua Lin | Hamid Alinejad-Rokny | Min Yang | Shiwen Ni | Wenhao Huang | Ge Zhang
Findings of the Association for Computational Linguistics: ACL 2025
Multimodal Large Language Models (MLLMs) are measured on numerous benchmarks like image captioning, visual question answer, and reasoning. However, these benchmarks often include overly simple or uninformative samples, making it difficult to effectively distinguish the performance of different MLLMs. Additionally, evaluating models across many benchmarks creates a significant computational burden. To address these issues, we propose LIME (Less Is More for MLLM Evaluation), a refined and efficient benchmark curated using a semi-automated pipeline. This pipeline filters out uninformative samples and eliminates answer leakage by focusing on tasks that require image-based understanding. Our experiments show that LIME reduces the number of samples by 76% and evaluation time by 77%, while it can more effectively distinguish different modelsā abilities. Notably, we find that traditional automatic metrics like CIDEr are insufficient for evaluating MLLMsā captioning performance, and excluding the caption task score yields a more accurate reflection of overall model performance. All code and data are available at https://anonymous.4open.science/r/LIME-49CD
AgentCourt: Simulating Court with Adversarial Evolvable Lawyer Agents
Guhong Chen | Liyang Fan | Zihan Gong | Nan Xie | Zixuan Li | Ziqiang Liu | Chengming Li | Qiang Qu | Hamid Alinejad-Rokny | Shiwen Ni | Min Yang
Findings of the Association for Computational Linguistics: ACL 2025
Guhong Chen | Liyang Fan | Zihan Gong | Nan Xie | Zixuan Li | Ziqiang Liu | Chengming Li | Qiang Qu | Hamid Alinejad-Rokny | Shiwen Ni | Min Yang
Findings of the Association for Computational Linguistics: ACL 2025
Current research in LLM-based simulation systems lacks comprehensive solutions for modeling real-world court proceedings, while existing legal language models struggle with dynamic courtroom interactions. We present AgentCourt, a comprehensive legal simulation framework that addresses these challenges through adversarial evolution of LLM-based agents. Our AgentCourt introduces a new adversarial evolutionary approach for agents called AdvEvol, which performs dynamic knowledge learning and evolution through structured adversarial interactions in a simulated courtroom program, breaking the limitations of the traditional reliance on static knowledge bases or manual annotations. By simulating 1,000 civil cases, we construct an evolving knowledge base that enhances the agentsā legal reasoning abilities. The evolved lawyer agents demonstrated outstanding performance on our newly introduced CourtBench benchmark, achieving a 12.1% improvement in performance compared to the original lawyer agents. Evaluations by professional lawyers confirm the effectiveness of our approach across three critical dimensions: cognitive agility, professional knowledge, and logical rigor. Beyond outperforming specialized legal models in interactive reasoning tasks, our findings emphasize the importance of adversarial learning in legal AI and suggest promising directions for extending simulation-based legal reasoning to broader judicial and regulatory contexts.
PEToolLLM: Towards Personalized Tool Learning in Large Language Models
Qiancheng Xu | Yongqi Li | Heming Xia | Fan Liu | Min Yang | Wenjie Li
Findings of the Association for Computational Linguistics: ACL 2025
Qiancheng Xu | Yongqi Li | Heming Xia | Fan Liu | Min Yang | Wenjie Li
Findings of the Association for Computational Linguistics: ACL 2025
Tool learning has emerged as a promising direction by extending Large Language Modelsā (LLMs) capabilities with external tools. Existing tool learning studies primarily focus on the general-purpose tool-use capability, which addresses explicit user requirements in instructions. However, they overlook the importance of personalized tool-use capability, leading to an inability to handle implicit user preferences. To address the limitation, we first formulate the task of personalized tool learning, which integrates userās interaction history towards personalized tool usage. To fill the gap of missing benchmarks, we construct PEToolBench, featuring diverse user preferences reflected in interaction history under three distinct personalized settings, and encompassing a wide range of tool-use scenarios. Moreover, we propose a framework PEToolLLaMA to adapt LLMs to the personalized tool learning task, which is trained through supervised fine-tuning and direct preference optimization. Extensive experiments on PEToolBench demonstrate the superiority of PEToolLLaMA over existing LLMs. We release code and data at https://github.com/travis-xu/PEToolBench.
STORYTELLER: An Enhanced Plot-Planning Framework for Coherent and Cohesive Story Generation
Jiaming Li | Yukun Chen | Ziqiang Liu | Minghuan Tan | Lei Zhang | Yunshui Li | Run Luo | Longze Chen | Jing Luo | Ahmadreza Argha | Hamid Alinejad-Rokny | Wei Zhou | Min Yang
Findings of the Association for Computational Linguistics: ACL 2025
Jiaming Li | Yukun Chen | Ziqiang Liu | Minghuan Tan | Lei Zhang | Yunshui Li | Run Luo | Longze Chen | Jing Luo | Ahmadreza Argha | Hamid Alinejad-Rokny | Wei Zhou | Min Yang
Findings of the Association for Computational Linguistics: ACL 2025
Stories are central to human culture, serving to share ideas, preserve traditions, and foster connections. Automatic story generation, a key advancement in artificial intelligence (AI), offers new possibilities for creating personalized content, exploring creative ideas, and enhancing interactive experiences. However, existing methods struggle to maintain narrative coherence and logical consistency. This disconnect compromises the overall storytelling experience, underscoring the need for substantial improvements. Inspired by human cognitive processes, we introduce Storyteller, a novel approach that systemically improves the coherence and consistency of automatically generated stories. Storyteller introduces a plot node structure based on linguistically grounded subject-verb-object (SVO) triplets, which capture essential story events and ensure a consistent logical flow. Unlike previous methods, Storyteller integrates two dynamic modulesāthe STORYLINE and narrative entity knowledge graph (NEKG)āthat continuously interact with the story generation process. This integration produces structurally sound, cohesive and immersive narratives. Extensive experiments demonstrate that Storyteller significantly outperforms existing approaches, achieving an 84.33% average win rate through human preference evaluation. At the same time, it is also far ahead in other aspects including creativity, coherence, engagement, and relevance.
Data Interpreter: An LLM Agent for Data Science
Sirui Hong | Yizhang Lin | Bang Liu | Bangbang Liu | Binhao Wu | Ceyao Zhang | Danyang Li | Jiaqi Chen | Jiayi Zhang | Jinlin Wang | Li Zhang | Lingyao Zhang | Min Yang | Mingchen Zhuge | Taicheng Guo | Tuo Zhou | Wei Tao | Robert Tang | Xiangtao Lu | Xiawu Zheng | Xinbing Liang | Yaying Fei | Yuheng Cheng | Yongxin Ni | Zhibin Gou | Zongze Xu | Yuyu Luo | Chenglin Wu
Findings of the Association for Computational Linguistics: ACL 2025
Sirui Hong | Yizhang Lin | Bang Liu | Bangbang Liu | Binhao Wu | Ceyao Zhang | Danyang Li | Jiaqi Chen | Jiayi Zhang | Jinlin Wang | Li Zhang | Lingyao Zhang | Min Yang | Mingchen Zhuge | Taicheng Guo | Tuo Zhou | Wei Tao | Robert Tang | Xiangtao Lu | Xiawu Zheng | Xinbing Liang | Yaying Fei | Yuheng Cheng | Yongxin Ni | Zhibin Gou | Zongze Xu | Yuyu Luo | Chenglin Wu
Findings of the Association for Computational Linguistics: ACL 2025
Large Language Model (LLM)-based agents have excelled in various domains but face significant challenges when applied to data science workflows due to their complex, multi-stage nature. Current LLM-based agents struggle with non-linear relationships, recursive dependencies, implicit data- and logic-dependent reasoning, and managing extensive context. In this paper, we introduce Data Interpreter, an LLM-based agent that addresses these challenges through hierarchical graph-based modeling to represent the complexity and a progressive strategy for step-by-step verification, refinement, and consistent context management. Extensive experiments confirm the effectiveness of Data Interpreter. On InfiAgent-DABench, it boosts performance by 25% (from 75.9% to 94.9%), and on machine learning and open-ended tasks, it lifts accuracy from 88% to 95% and from 60% to 97%, respectively. Moreover, our method surpasses state-of-the-art baselines by 26% on the MATH dataset. We will release the code upon publication.
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct
Run Luo | Haonan Zhang | Longze Chen | Ting-En Lin | Xiong Liu | Yuchuan Wu | Min Yang | Yongbin Li | Minzheng Wang | Pengpeng Zeng | Lianli Gao | Heng Tao Shen | Yunshui Li | Hamid Alinejad-Rokny | Xiaobo Xia | Jingkuan Song | Fei Huang
Findings of the Association for Computational Linguistics: ACL 2025
Run Luo | Haonan Zhang | Longze Chen | Ting-En Lin | Xiong Liu | Yuchuan Wu | Min Yang | Yongbin Li | Minzheng Wang | Pengpeng Zeng | Lianli Gao | Heng Tao Shen | Yunshui Li | Hamid Alinejad-Rokny | Xiaobo Xia | Jingkuan Song | Fei Huang
Findings of the Association for Computational Linguistics: ACL 2025
The development of Multimodal Large Language Models (MLLMs) has seen significant progress, driven by increasing demands across various fields (e.g., multimodal agents, embodied intelligence). While model-driven approaches aim to enhance MLLM capabilities through diverse architectures, their performance gains have become increasingly marginal. In contrast, data-driven methods, which scale up image-text instruction datasets, have proven more effective but face challenges related to limited data diversity and complexity. The absence of high-quality instruction data remains a major bottleneck in MLLM development. To address this issue, we propose , a novel multimodal instruction data evolution framework. This framework iteratively enhances data quality through a refined combination of fine-grained perception, cognitive reasoning, and interaction evolution, generating a more complex and diverse image-text instruction dataset that significantly improves MLLM capabilities. Starting with an initial dataset, SEED-163K, we employ to systematically expand instruction diversity, extend visual reasoning steps to improve cognitive abilities, and extract fine-grained visual details to enhance understanding and robustness. To rigorously evaluate our approach, we conduct extensive qualitative analysis and quantitative experiments across 13 vision-language tasks. Compared to baseline models trained on the original seed dataset, our method achieves an average accuracy improvement of 3.1 percentage points. Moreover, our approach attains state-of-the-art (SOTA) performance in nine tasks while using significantly less data than existing state-of-the-art models.
Expanding before Inferring: Enhancing Factuality in Large Language Models through Premature Layers Interpolation
Dingwei Chen | Ziqiang Liu | Feiteng Fang | Chak Tou Leong | Shiwen Ni | Ahmadreza Argha | Hamid Alinejad-Rokny | Min Yang | Chengming Li
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Dingwei Chen | Ziqiang Liu | Feiteng Fang | Chak Tou Leong | Shiwen Ni | Ahmadreza Argha | Hamid Alinejad-Rokny | Min Yang | Chengming Li
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Large Language Models (LLMs) demonstrate remarkable capabilities in text understanding and generation. However, their tendency to produce factually inconsistent outputsācommonly referred to as āhallucinationsāāremains a critical challenge. Existing approaches, such as retrieval-based and inference-time correction methods, primarily address this issue at the input or output level, often overlooking the intrinsic information refinement process and the role of premature layers. Meanwhile, alignment- and fine-tuning-based methods are resource-intensive. In this paper, we propose PLI (Premature Layers Interpolation), a novel, training-free, and plug-and-play intervention designed to enhance factuality. PLI mitigates hallucinations by inserting premature layers formed through mathematical interpolation with adjacent layers. Inspired by stable diffusion and sampling steps, PLI extends the depth of information processing and transmission in LLMs, improving factual coherence. Experiments on four publicly available datasets demonstrate that PLI effectively reduces hallucinations while outperforming existing baselines in most cases. Further analysis suggests that the success of layer interpolation is closely linked to LLMsā internal mechanisms. To promote reproducibility, we will release our code and data upon acceptance.
UNComp: Can Matrix Entropy Uncover Sparsity? ā A Compressor Design from an Uncertainty-Aware Perspective
Jing Xiong | Jianghan Shen | Fanghua Ye | Chaofan Tao | Zhongwei Wan | Jianqiao Lu | Xun Wu | Chuanyang Zheng | Zhijiang Guo | Min Yang | Lingpeng Kong | Ngai Wong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jing Xiong | Jianghan Shen | Fanghua Ye | Chaofan Tao | Zhongwei Wan | Jianqiao Lu | Xun Wu | Chuanyang Zheng | Zhijiang Guo | Min Yang | Lingpeng Kong | Ngai Wong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Deploying large language models (LLMs) for long-context inference remains challenging due to their substantial memory and computational demands. While techniques such as Key-Value (KV) cache compression are designed to reduce memory usage, they often neglect the structured sparsity inherent in the relationship between hidden states and their corresponding KV cache. In this work, we explore the role of uncertainty as a potential indicator of sparsity within LLMs. We propose UNComp, an uncertainty-aware framework that leverages truncated matrix entropy to identify areas of low information content, thereby revealing sparsity patterns that can be used for adaptive compression. Unlike traditional methods that apply uniform compression, UNComp dynamically adjusts its approach to compression, guided by uncertainty measures that reflect the importance of various model components. Our analysis shows that sparsity patterns, when derived from uncertainty estimates, can be exploited to reveal special long-range dependencies, such as retrieval heads and retrieval layers. This perspective not only enhances our understanding of how compression can be optimized but also provides new insights into the inherent sparsity of LLMs during long-context inference. By focusing on uncertainty to analyze the sparsity pattern in detail, UNComp reduces the KV cache size to 4.74% of the original, achieves a 6% prefill speedup, and improves throughput by 6.4Ć ā not only delivering strong lossless compression performance, but also validating the effectiveness of the underlying theoretical tool. Our codes are submitted with the paper.
Exploring the Impact of Personality Traits on LLM Bias and Toxicity
Shuo Wang | Renhao Li | Xi Chen | Yulin Yuan | Min Yang | Derek F. Wong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Shuo Wang | Renhao Li | Xi Chen | Yulin Yuan | Min Yang | Derek F. Wong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
With the different roles that AI is expected to play in human life, imbuing large language models (LLMs) with different personalities has attracted increasing research interest. While the āpersonificationā enhances human experiences of interactivity and adaptability of LLMs, it gives rise to critical concerns about content safety, particularly regarding bias, sentiment, and toxicity of LLM generation. This study explores how assigning different personality traits to LLMs affects the toxicity and biases of their outputs. Leveraging the widely accepted HEXACO personality framework developed in social psychology, we design experimentally sound prompts to test three LLMsā performance on three toxic and bias benchmarks. The findings demonstrate the sensitivity of all three models to HEXACO personality traits and, more importantly, a consistent variation in the biases, negative sentiment, and toxicity of their output. In particular, adjusting the levels of several personality traits can effectively reduce bias and toxicity in model performance, similar to humansā correlations between personality traits and toxic behaviors. The findings highlight the additional need to examine content safety besides the efficiency of training or fine-tuning methods for LLM personification, they also suggest a potential for the adjustment of personalities to be a simple and low-cost method to conduct controlled text generation.
Can MLLMs Understand the Deep Implication Behind Chinese Images?
Chenhao Zhang | Xi Feng | Yuelin Bai | Xeron Du | Jinchang Hou | Kaixin Deng | Guangzeng Han | Qinrui Li | Bingli Wang | Jiaheng Liu | Xingwei Qu | Yifei Zhang | Qixuan Zhao | Yiming Liang | Ziqiang Liu | Feiteng Fang | Min Yang | Wenhao Huang | Chenghua Lin | Ge Zhang | Shiwen Ni
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chenhao Zhang | Xi Feng | Yuelin Bai | Xeron Du | Jinchang Hou | Kaixin Deng | Guangzeng Han | Qinrui Li | Bingli Wang | Jiaheng Liu | Xingwei Qu | Yifei Zhang | Qixuan Zhao | Yiming Liang | Ziqiang Liu | Feiteng Fang | Min Yang | Wenhao Huang | Chenghua Lin | Ge Zhang | Shiwen Ni
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
As the capabilities of Multimodal Large Language Models (MLLMs) improve, the need for higher-order evaluation of them is increasing. However, there is a lack of work evaluating MLLM for higher-order perception and understanding of Chinese visual content. To address this, we introduce the CII-Bench, which aims to assess MLLMsā such capabilities for Chinese images. To ensure the authenticity of the Chinese context, images in CII-Bench are sourced from the Chinese Internet and manually reviewed, with corresponding answers also manually crafted. Additionally, CII-Bench incorporates images that represent Chinese traditional culture, such as famous Chinese traditional paintings, which can deeply reflect the modelās understanding of Chinese traditional culture. Through experiments on multiple MLLMs using CII-Bench, significant findings emerged. There is a large gap between MLLMs and humans in performance. The highest MLLM accuracy is 64.4%, while the human average is 78.2% and the peak is 81.0%. MLLMs perform poorly on traditional culture images, indicating limitations in understanding high-level semantics and lacking a deep knowledge base of Chinese traditional culture. Moreover, most models have higher accuracy when image emotion hints are added to the prompts. We believe CII-Bench will help MLLMs better understand Chinese semantics and specific images, and move forward the development of expert artificial general intelligence (AGI). Our project is publicly available at https://cii-bench.github.io.
Learning First-Order Logic Rules for Argumentation Mining
Yang Sun | Guanrong Chen | Hamid Alinejad-Rokny | Jianzhu Bao | Yuqi Huang | Bin Liang | Kam-Fai Wong | Min Yang | Ruifeng Xu
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yang Sun | Guanrong Chen | Hamid Alinejad-Rokny | Jianzhu Bao | Yuqi Huang | Bin Liang | Kam-Fai Wong | Min Yang | Ruifeng Xu
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Argumentation Mining (AM) aims to extract argumentative structures from texts by identifying argumentation components (ACs) and their argumentative relations (ARs). While previous works focus on representation learning to encode ACs and AC pairs, they fail to explicitly model the underlying reasoning patterns of AM, resulting in limited interpretability. This paper proposes a novel  ̲First- ̲Order  ̲Logic reasoning framework for  ̲AM (FOL-AM), designed to explicitly capture logical reasoning paths within argumentative texts. By interpreting multiple AM subtasks as a unified relation query task modeled using FOL rules, FOL-AM facilitates multi-hop relational reasoning and enhances interpretability. The framework supports two flexible implementations: a fine-tuned approach to leverage task-specific learning, and a prompt-based method utilizing large language models to harness their generalization capabilities. Extensive experiments on two AM benchmarks demonstrate that FOL-AM outperforms strong baselines while significantly improving explainability.
Quantification of Large Language Model Distillation
Sunbowen Lee | Junting Zhou | Chang Ao | Kaige Li | Xeron Du | Sirui He | Haihong Wu | Tianci Liu | Jiaheng Liu | Hamid Alinejad-Rokny | Min Yang | Yitao Liang | Zhoufutu Wen | Shiwen Ni
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Sunbowen Lee | Junting Zhou | Chang Ao | Kaige Li | Xeron Du | Sirui He | Haihong Wu | Tianci Liu | Jiaheng Liu | Hamid Alinejad-Rokny | Min Yang | Yitao Liang | Zhoufutu Wen | Shiwen Ni
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Model distillation is a fundamental technique in building large language models (LLMs), transferring knowledge from a teacher model to a student model. However, distillation can lead to model homogenization, reducing diversity among models and impairing their ability to robustly handle complex or novel tasks. These limitations underscore the need to systematically quantify the distillation process and its impact. In this work, we propose a framework to evaluate and quantify model distillation. Our method addresses two key aspects: (1) Identifying identity cognition contradictions to assess discrepancies in how models perceive and represent identity-related information, and (2) Analyzing multi-granularity response similarities across models to measure the extent of homogenization. Experimental results demonstrate two key insights: (1) Well-known closed-source and open-source LLMs usually exhibit high distillation degrees, except for Claude, Doubao, and Gemini. (2) Base LLMs show higher distillation degrees compared to aligned LLMs. By offering a systematic approach to improve the transparency of LLM data distillation, we call for LLMs with more independent development and more transparent technical reports to improve LLMsā robustness and safety. The code and data are available at https://github.com/Aegis1863/LLMs-Distillation-Quantification.
CLaSp: In-Context Layer Skip for Self-Speculative Decoding
Longze Chen | Renke Shan | Huiming Wang | Lu Wang | Ziqiang Liu | Run Luo | Jiawei Wang | Hamid Alinejad-Rokny | Min Yang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Longze Chen | Renke Shan | Huiming Wang | Lu Wang | Ziqiang Liu | Run Luo | Jiawei Wang | Hamid Alinejad-Rokny | Min Yang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Speculative decoding (SD) is a promising method for accelerating the decoding process of Large Language Models (LLMs). The efficiency of SD primarily hinges on the consistency between the draft model and the verify model. However, existing drafting approaches typically require additional modules to be trained, which can be challenging to implement and ensure compatibility across various LLMs. In this paper, we propose CLaSp, an in-context layer-skipping strategy for self-speculative decoding. Unlike prior methods, CLaSp does not require additional drafting modules or extra training. Instead, it employs a plug-and-play mechanism by skipping intermediate layers of the verify model to construct a compressed draft model. Specifically, we develop a dynamic programming algorithm that optimizes the layer-skipping process by leveraging the complete hidden states from the last verification stage as an objective. This enables CLaSp to dynamically adjust its layer-skipping strategy after each verification stage, without relying on pre-optimized sets of skipped layers. Experimental results across diverse downstream tasks demonstrate that CLaSp achieves a speedup of 1.3Ć ā¼ 1.7Ć on LLaMA3 series models without altering the original distribution of the generated text.
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
Haonan Zhang | Run Luo | Xiong Liu | Yuchuan Wu | Ting-En Lin | Pengpeng Zeng | Qiang Qu | Feiteng Fang | Min Yang | Lianli Gao | Jingkuan Song | Fei Huang | Yongbin Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Haonan Zhang | Run Luo | Xiong Liu | Yuchuan Wu | Ting-En Lin | Pengpeng Zeng | Qiang Qu | Feiteng Fang | Min Yang | Lianli Gao | Jingkuan Song | Fei Huang | Yongbin Li
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Role-Playing Agents (RPAs), benefiting from large language models, is an emerging interactive AI system that simulates roles or characters with diverse personalities. However, existing methods primarily focus on mimicking dialogues among roles in textual form, neglecting the roleās voice traits (e.g., voice style and emotions) as playing a crucial effect in interaction, which tends to be more immersive experiences in realistic scenarios. Towards this goal, we propose OmniCharacter, a first seamless speech-language personality interaction model to achieve immersive RPAs with low latency. Specifically, OmniCharacter enables agents to consistently exhibit role-specific personality traits and vocal traits throughout the interaction, enabling a mixture of speech and language responses. To align the model with speech-language scenarios, we construct a dataset named OmniCharacter-10K, which involves more distinctive characters (20), richly contextualized multi-round dialogue (10K), and dynamic speech response (135K). Experimental results showcase that our method yields better responses in terms of both content and style compared to existing RPAs and mainstream speech-language models, with a response latency as low as 289ms.
Search
Fix author
Co-authors
- Hamid Alinejad-Rokny 14
- Shiwen Ni 10
- Feiteng Fang 8
- Ziqiang Liu 8
- Run Luo 5
- Guhong Chen 4
- Longze Chen 4
- Wenhao Huang 4
- Chengming Li 4
- Yongbin Li 4
- Jiaheng Liu 4
- Ge Zhang 4
- Ahmadreza Argha 3
- Xeron Du 3
- Liyang Fan 3
- Renhao Li 3
- Chenghua Lin 3
- Ting-En Lin 3
- Xiong Liu 3
- Qiang Qu 3
- Minghuan Tan 3
- Derek F. Wong (é»č¾) 3
- Yuchuan Wu 3
- Haonan Zhang 3
- Yuelin Bai 2
- Dingwei Chen 2
- Lianli Gao 2
- Zihan Gong 2
- Fei Huang 2
- Feng Jiang (čå³°) 2
- Jiaming Li 2
- Shuaimin Li 2
- Wenjie Li 2
- Yizhi Li 2
- Yongqi Li 2
- Yunshui Li 2
- Yiming Liang 2
- Fan Liu 2
- Xingwei Qu 2
- Jingkuan Song 2
- Qiyao Wang 2
- Zhoufutu Wen 2
- Haihong Wu 2
- Qiancheng Xu 2
- Pengpeng Zeng 2
- Lei Zhang 2
- Tianyu Zheng 2
- Junting Zhou 2
- Chang Ao 1
- Lu Bai 1
- Jianzhu Bao 1
- Mingshan Chang 1
- Guangxu Chen 1
- Guanrong Chen 1
- Hardy Chen 1
- Jiaqi Chen 1
- Xi Chen 1
- Xin Chen 1
- Yukun Chen 1
- Yuheng Cheng 1
- Lixin Cui 1
- Kaixin Deng 1
- Yaying Fei 1
- Xi Feng 1
- Cheng Fu 1
- Jie Fu 1
- Zhibin Gou 1
- Jiawei Guo 1
- Shuyue Guo 1
- Taicheng Guo 1
- Zhijiang Guo 1
- Guangzeng Han 1
- Qi Han 1
- Sirui He 1
- Sirui Hong 1
- Jinchang Hou 1
- Xiping Hu 1
- Fei Huang 1
- Ruijun Huang 1
- Shujian Huang (书å é») 1
- Xiang Huang 1
- Yuqi Huang 1
- Zhihong Huang 1
- Shian Jia 1
- Leo Jin 1
- Kun Jing 1
- Zhiqiao Kang 1
- Philipp Koehn 1
- Lingpeng Kong 1
- Sunbowen Lee 1
- Chak Tou Leong 1
- Binhua Li 1
- Bo Li 1
- Danyang Li 1
- Jiayan Li 1
- Junxiong Li 1
- Kaige Li 1
- Qinrui Li 1
- Siyi Li 1
- Yanhong Li 1
- Zeyang Li 1
- Zixuan Li 1
- Bin Liang (ę¢ę) 1
- Xinbing Liang 1
- Yitao Liang 1
- Hongquan Lin 1
- Junyang Lin 1
- Yizhang Lin 1
- Yuan Lin 1
- Yufang Lin 1
- Li Linwei 1
- Bang Liu 1
- Bangbang Liu 1
- Huaren Liu 1
- Ruibo Liu 1
- Tianci Liu 1
- Yantao Liu 1
- Jianqiao Lu 1
- Xiangtao Lu 1
- Jing Luo 1
- Yuyu Luo 1
- Ziyu Lyu 1
- Nuo Ma 1
- Yongxin Ni 1
- Yupeng Qi 1
- Zhifei Qin 1
- Renke Shan 1
- Heng Tao Shen 1
- Jianghan Shen 1
- Jiajun Shi 1
- Yuanfeng Song 1
- Yang Su 1
- Chenghao Sun 1
- Yang Sun 1
- Robert Tang 1
- Chaofan Tao 1
- Wei Tao 1
- Jianhong Tu 1
- Zhongwei Wan 1
- Zhuoyue Wan 1
- Bingli Wang 1
- Hongbo Wang 1
- Hongru Wang 1
- Huiming Wang 1
- Jiawei Wang 1
- Jinlin Wang 1
- Lu Wang 1
- Minzheng Wang 1
- Shijian Wang 1
- Shiqiang Wang 1
- Shuo Wang 1
- Songsheng Wang 1
- Yanlin Wang 1
- Zekun Moore Wang 1
- ChaoPeng Wei 1
- HU Wei 1
- Kam-Fai Wong 1
- Ngai Wong 1
- Binhao Wu 1
- Chenglin Wu 1
- Haoning Wu 1
- Siwei Wu 1
- Xun Wu 1
- Heming Xia 1
- Xiaobo Xia 1
- Nan Xie 1
- Wenya Xie 1
- Jing Xiong 1
- Kan Xu 1
- Ruifeng Xu (å¾ēæå³°) 1
- Xander Xu 1
- Zongze Xu 1
- Shuo Yan 1
- Jian Yang 1
- Fanghua Ye 1
- Jing Ye 1
- Wu Yihang 1
- Tianshu Yu 1
- Ruibin Yuan 1
- Yulin Yuan 1
- Xiang Yue 1
- Yilin Yue 1
- Qianbo Zang 1
- Ceyao Zhang 1
- Chen Jason Zhang 1
- Chenhao Zhang 1
- Jiajun Zhang 1
- Jiayi Zhang 1
- Li Zhang 1
- Lingyao Zhang 1
- Shijie Zhang 1
- Xincheng Zhang 1
- Yifei Zhang 1
- Yiqian Zhang 1
- Bing Zhao 1
- Jiahao Zhao 1
- Qixuan Zhao 1
- Chuanyang Zheng 1
- Xiawu Zheng 1
- Tuo Zhou 1
- Wei Zhou 1
- King Zhu 1
- Liang Zhu 1
- Minghui Zhu 1
- Yuxuan Zhu 1
- Mingchen Zhuge 1