Yue Huang
Author directoryPapers on this page may belong to the following people: Yue Huang, Yue Huang, Yue Huang
2026
Automatic Prompt Engineering for Generative AI–Based Essay Scoring
Yue Huang
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Yue Huang
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study evaluated automatic prompt engineering (APE) using one assignment in the PERSUADE 2.0 dataset. The APE approach achieved higher QWK (.812) than the research-informed, zero-shot baseline prompting approach (.646). Descriptive comparisons examined gender and English language learner subgroups. Findings support the benefits of APE for automated essay scoring (AES).
Rubric-Aligned Generative-AI Features as Supplementary Predictors in Feature-Based Automated Essay Scoring
Yue Huang | Duanli Yan | Corey Palermo
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
Yue Huang | Duanli Yan | Corey Palermo
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Coordinated Session Papers
This study examined whether rubric-aligned generative-AI features could augment established linguistic features in trait-based automated essay scoring. Features from both sources showed meaningful associations with human scores and only partial overlap with one another. Scoring models combining both feature sets produced modest improvements that varied across traits and evaluation metrics.
Prompting as Measurement: Enhancing Coding for Systematic Reviews with Large Language Models
Dandan Chen Kaptur | Yue Huang | Yanhui Guo | Xuejun Ryan Ji
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Dandan Chen Kaptur | Yue Huang | Yanhui Guo | Xuejun Ryan Ji
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers
Prompting functions as a measurement condition that alters LLM coding behavior. We examined whether few-shot prompting improves GPT-based qualitative coding in systematic reviews. Contrary to expectations, we found that the zero-shot approach produced the highest agreement with human coders, suggesting increasing conservatism as more prompt structure was added.
2025
TRUSTEVAL: A Dynamic Evaluation Toolkit on Trustworthiness of Generative Foundation Models
Yanbo Wang | Jiayi Ye | Siyuan Wu | Chujie Gao | Yue Huang | Xiuying Chen | Yue Zhao | Xiangliang Zhang
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)
Yanbo Wang | Jiayi Ye | Siyuan Wu | Chujie Gao | Yue Huang | Xiuying Chen | Yue Zhao | Xiangliang Zhang
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations)
Ensuring the trustworthiness of Generative Foundation Models (GenFMs) is a pressing challenge as they gain widespread use. Existing evaluation toolkits are often limited in scope, dynamism, and flexibility. This paper introduces TRUSTEVAL, a dynamic and comprehensive toolkit designed for evaluating GenFMs across various dimensions. TRUSTEVAL supports both dynamic dataset generation and evaluation, offering advanced features including comprehensiveness, usability, and flexibility. TRUSTEVAL integrates diverse generative models, datasets, evaluation methods, metrics, inference efficiency enhancement, and evaluation report generation. Through case studies, we demonstrate TRUSTEVAL’s potential to advance the trustworthiness evaluation of GenFMs.
Evaluating LLM-Based Automated Essay Scoring: Accuracy, Fairness, and Validity
Yue Huang | Joshua Wilson
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Yue Huang | Joshua Wilson
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study evaluates large language models (LLMs) for automated essay scoring (AES), comparing prompt strategies and fairness across student groups. We found that well-designed prompting helps LLMs approach traditional AES performance, but both differ from human scores for ELLs—the traditional model shows larger overrall gaps, while LLMs show subtler disparities.
Comparison of AI and Human Scoring on A Visual Arts Assessment
Ning Jiang | Yue Huang | Jie Chen
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Ning Jiang | Yue Huang | Jie Chen
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
This study examines reliability and comparability of Generative AI scores versus human ratings on two performance tasks—text-based and drawing-based—in a fourth-grade visual arts assessment. Results show GPT-4 is consistent, aligned with humans but more lenient, and its agreement with humans is slightly lower than that between human raters.
2024
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected?
Qihui Zhang | Chujie Gao | Dongping Chen | Yue Huang | Yixin Huang | Zhenyang Sun | Shilin Zhang | Weiye Li | Zhengyan Fu | Yao Wan | Lichao Sun
Findings of the Association for Computational Linguistics: NAACL 2024
Qihui Zhang | Chujie Gao | Dongping Chen | Yue Huang | Yixin Huang | Zhenyang Sun | Shilin Zhang | Weiye Li | Zhengyan Fu | Yao Wan | Lichao Sun
Findings of the Association for Computational Linguistics: NAACL 2024
With the rapid development and widespread application of Large Language Models (LLMs), the use of Machine-Generated Text (MGT) has become increasingly common, bringing with it potential risks, especially in terms of quality and integrity in fields like news, education, and science. Current research mainly focuses on purely MGT detection, without adequately addressing mixed scenarios including AI-revised Human-Written Text (HWT) or human-revised MGT. To tackle this challenge, we define mixtext, a form of mixed text involving both AI and human-generated content. Then we introduce MixSet, the first dataset dedicated to studying these mixtext scenarios. Leveraging MixSet, we executed comprehensive experiments to assess the efficacy of prevalent MGT detectors in handling mixtext situations, evaluating their performance in terms of effectiveness, robustness, and generalization. Our findings reveal that existing detectors struggle to identify mixtext, particularly in dealing with subtle modifications and style adaptability. This research underscores the urgent need for more fine-grain detectors tailored for mixtext, offering valuable insights for future research. Code and Models are available at https://github.com/Dongping-Chen/MixSet.
AlignBench: Benchmarking Chinese Alignment of Large Language Models
Xiao Liu | Xuanyu Lei | Shengyuan Wang | Yue Huang | Andrew Feng | Bosi Wen | Jiale Cheng | Pei Ke | Yifan Xu | Weng Lam Tam | Xiaohan Zhang | Lichao Sun | Xiaotao Gu | Hongning Wang | Jing Zhang | Minlie Huang | Yuxiao Dong | Jie Tang
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xiao Liu | Xuanyu Lei | Shengyuan Wang | Yue Huang | Andrew Feng | Bosi Wen | Jiale Cheng | Pei Ke | Yifan Xu | Weng Lam Tam | Xiaohan Zhang | Lichao Sun | Xiaotao Gu | Hongning Wang | Jing Zhang | Minlie Huang | Yuxiao Dong | Jie Tang
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Alignment has become a critical step for instruction-tuned Large Language Models (LLMs) to become helpful assistants. However, effective evaluation of alignment for emerging Chinese LLMs is still significantly lacking, calling for real-scenario grounded, open-ended, challenging and automatic evaluations tailored for alignment. To fill in this gap, we introduce AlignBench, a comprehensive multi-dimensional benchmark for evaluating LLMs’ alignment in Chinese. We tailor a human-in-the-loop data curation pipeline, containing 8 main categories, 683 real-scenario rooted queries and corresponding human verified references.To ensure references’ correctness, each knowledge-intensive query is accompanied with evidences collected from reliable webpages (including the url and quotation) by our annotators.For automatic evaluation, our benchmark employs a rule-calibrated multi-dimensional LLM-as-Judge (CITATION) with Chain-of-Thought to generate explanations and final ratings as evaluations, ensuring high reliability and interpretability.All evaluation codes and data are publicly available at https://github.com/THUDM/AlignBench
Search
Fix author
Co-authors
- Chujie Gao 2
- Lichao Sun 2
- Dongping Chen 1
- Jie Chen 1
- Xiuying Chen 1
- Jiale Cheng 1
- Yuxiao Dong 1
- Andrew Feng 1
- Zhengyan Fu 1
- Xiaotao Gu 1
- Yanhui Guo 1
- Minlie Huang 1
- Yixin Huang 1
- Xuejun Ryan Ji 1
- Ning Jiang 1
- Dandan Chen Kaptur 1
- Pei Ke 1
- Xuanyu Lei 1
- Weiye Li 1
- Xiao Liu 1
- Corey Palermo 1
- Zhenyang Sun 1
- Weng Lam Tam 1
- Jie Tang 1
- Yao Wan 1
- Hongning Wang 1
- Shengyuan Wang 1
- Yanbo Wang 1
- Bosi Wen 1
- Joshua Wilson 1
- Siyuan Wu 1
- Yifan Xu 1
- Duanli Yan 1
- Jiayi Ye 1
- Jing Zhang 1
- Qihui Zhang 1
- Shilin Zhang 1
- Xiangliang Zhang 1
- Xiaohan Zhang 1
- Yue Zhao 1