Minggui He
Author directoryAlso published as: Minggui HE
2026
R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning
Minggui He | Yilun Liu | Shimin Tao | Hongyong Zeng | Jian Zhang | Yuanchang Luo | Li Zhang | Daimeng Wei | Weibin Meng | Osamu Yoshie
Transactions of the Association for Computational Linguistics, Volume 14
Minggui He | Yilun Liu | Shimin Tao | Hongyong Zeng | Jian Zhang | Yuanchang Luo | Li Zhang | Daimeng Wei | Weibin Meng | Osamu Yoshie
Transactions of the Association for Computational Linguistics, Volume 14
Despite recent breakthroughs in reasoning-enhanced large language models (LLMs), incorporating inference-time reasoning into application tasks such as machine translation (MT), where human translators naturally employ structured, multi-layered reasoning chain-of-thoughts (CoTs), is yet un-derexplored. Existing methods either design a fixed CoT tailored for a specific MT sub-task (e.g., literature translation), or rely on synthesizing CoTs unaligned with humans and supervised fine-tuning (SFT) prone to overfitting, limiting their adaptability to diverse translation scenarios. This paper introduces R1-Translator (R1-T1), a novel framework to achieve inference-time reasoning for general MT via reinforcement learning (RL) with human-aligned CoTs comprising six common patterns. Our approach pioneers three innovations: (1) verifying reasoning-based translation in various MT scenarios (e.g., multilingual MT, domain MT) unseen from the training phase; (2) formalizing six expert-curated CoT templates that mirror hybrid human strategies like context-aware paraphrasing and round-trip translation; and (3) enabling more flexible CoTs through an RL stage after cold-start. Both human and automatic evaluation results indicate a steady translation quality improvement in a total of 10+ languages and 40+ translation directions on Flores-101 test set and four domain-specific MT tasks, especially on the languages unseen from training.
The GaoYao Benchmark: A Comprehensive Framework for Evaluating Multilingual and Multicultural Abilities of Large Language Models
Yilun Liu | Chunguang Zhao | Mengyao Piao | Lingqi Miao | Shimin Tao | Minggui HE | Chenxin Liu | Zhang Li | Mahongxia | Jiaxin Guo | Chen Liu | Liqun Deng | Jiansheng Wei | Xiaojun Meng | Fanyi Du | Daimeng Wei | Yanghua Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yilun Liu | Chunguang Zhao | Mengyao Piao | Lingqi Miao | Shimin Tao | Minggui HE | Chenxin Liu | Zhang Li | Mahongxia | Jiaxin Guo | Chen Liu | Liqun Deng | Jiansheng Wei | Xiaojun Meng | Fanyi Du | Daimeng Wei | Yanghua Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often neglect deep cultural nuances; (2) insufficient language coverage in subjective tasks relying on low-quality machine translation; and (3) shallow analysis that lacks diagnostic depth beyond simple rankings. To address these, we introduce GaoYao, a comprehensive benchmark with 182.3k samples, 26 languages and 51 nations/areas. First, GaoYao proposes a unified framework categorizing evaluation tasks into three cultural layers (General Multilingual, Cross-cultural, Monocultural) and nine cognitive sub-layers. Second, we achieve native-quality expansion by leveraging experts to rigorously localize subjective benchmarks into 19 languages and synthesizing cross-cultural test sets for 34 cultures, surpassing prior coverage by up to 111%. Third, we conduct an in-depth diagnostic analysis on 20+ flagship and compact LLMs. Our findings reveal significant geographical performance disparities and distinct gaps between tasks, offering a reliable map for future work. We release the benchmark.
The “Knowledge–Behavior Gap” in Cultural Taboo Safety of Large Language Models
Ying He | Sihang Jiang | Xingzhou Chen | Zhouhong Gu | Yiwei Gu | Minggui HE | Shimin Tao | Mahongxia | Yanghua Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Ying He | Sihang Jiang | Xingzhou Chen | Zhouhong Gu | Yiwei Gu | Minggui HE | Shimin Tao | Mahongxia | Yanghua Xiao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or values biases, while overlooking whether LLMs can recognize and respect cultural taboos, especially when taboos are implicitly hidden in seemingly harmless questions. Besides, cultural taboos are implicit, and context-dependent, thus poss unique challenges for reliable evaluation. To address these gaps, we introduce CulShield, the first public benchmark dedicated to evaluating and improving the cultural taboo safety of LLMs. CulShield spans 77 countries and regions, and includes over 2,020 taboos. It evaluates models along both explicit knowledge and implicit behaviors.Experiments on several advanced LLMs (e.g., GPT-4o-mini, Gemini-2.5-pro) reveal a clear “knowledge-behavior gap”: models often fail to apply known taboos during interaction. We further show that variations in linguistic context can significantly affect LLMs’ cultural taboo safety. Code and data is accessible here: https://anonymous.4open.science/r/CulShield-7A0E.
2025
Taming Text-to-Image Synthesis for Novices: User-centric Prompt Generation via Multi-turn Guidance
Yilun Liu | Minggui He | Feiyu Yao | Yuhe Ji | Shimin Tao | Jingzhou Du | Justin Li | Jian Gao | Zhang Li | Hao Yang | Boxing Chen | Osamu Yoshie
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Yilun Liu | Minggui He | Feiyu Yao | Yuhe Ji | Shimin Tao | Jingzhou Du | Justin Li | Jian Gao | Zhang Li | Hao Yang | Boxing Chen | Osamu Yoshie
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
The emergence of text-to-image synthesis (TIS) models has significantly influenced digital image creation by producing high-quality visuals from written descriptions. Yet these models are sensitive on textual prompts, posing a challenge for novice users who may not be familiar with TIS prompt writing. Existing solutions relieve this via automatic prompt expansion or generation from a user query. However, this single-turn manner suffers from limited user-centricity in terms of result interpretability and user interactivity. Thus, we propose DialPrompt, a dialogue-based TIS prompt generation model that emphasizes user experience for novice users. DialPrompt is designed to follow a multi-turn workflow, where in each round of dialogue the model guides user to express their preferences on possible optimization dimensions before generating the final TIS prompt. To achieve this, we mined 15 essential dimensions for high-quality prompts from advanced users and curated a multi-turn dataset. Through training on this dataset, DialPrompt improves user-centricity by allowing users to perceive and control the creation process of TIS prompts. Experiments indicate that DialPrompt improves significantly in user-centricity score compared with existing approaches while maintaining a competitive quality of synthesized images. In our user evaluation, DialPrompt is highly rated by 19 human reviewers (especially novices).
Search
Fix author
Co-authors
- Shimin Tao 4
- Zhang Li 2
- Yilun Liu 2
- Mahongxia 2
- Daimeng Wei 2
- Yanghua Xiao 2
- Osamu Yoshie 2
- Boxing Chen 1
- Xingzhou Chen 1
- Liqun Deng 1
- Fanyi Du 1
- Jingzhou Du 1
- Jian Gao 1
- Yiwei Gu 1
- Zhouhong Gu 1
- Jiaxin Guo 1
- Ying He 1
- Yuhe Ji 1
- Sihang Jiang 1
- Justin Li 1
- Chen Liu 1
- Chenxin Liu 1
- Yilun Liu 1
- Yuanchang Luo 1
- Weibin Meng 1
- Xiaojun Meng 1
- Lingqi Miao 1
- Mengyao Piao 1
- Jiansheng Wei 1
- Hao Yang 1
- Feiyu Yao 1
- Hongyong Zeng 1
- Jian Zhang 1
- Li Zhang 1
- Chunguang Zhao 1