Shuai Wang
Author directoryOther people with similar names: Shuai Wang, Shuai Wang, Shuai Wang, Shuai Wang, Shuai Wang
Unverified author pages with similar names: Shuai Wang
2026
PromptPrism: A Linguistically-Inspired Taxonomy for Prompts
Sullam Jeoung | Yueyan Chen | Yi Zhang | Shuai Wang | Haibo Ding | Lin Lee Cheong
Findings of the Association for Computational Linguistics: EACL 2026
Sullam Jeoung | Yueyan Chen | Yi Zhang | Shuai Wang | Haibo Ding | Lin Lee Cheong
Findings of the Association for Computational Linguistics: EACL 2026
Prompts are the interface for eliciting the capabilities of large language models (LLMs). Understanding their structure and components is critical for analyzing LLM behavior and optimizing performance. However, the field lacks a comprehensive framework for systematic prompt analysis and understanding. We introduce PromptPrism, a linguistically-inspired taxonomy that enables prompt analysis across three hierarchical levels: functional structure, semantic component, and syntactic pattern. By applying linguistic concepts to prompt analysis, PromptPrism bridges traditional language understanding and modern LLM research, offering insights that purely empirical approaches might miss. We show the practical utility of PromptPrism by applying it to three applications: (1) a taxonomy-guided prompt refinement approach that automatically improves prompt quality and enhances model performance across a range of tasks; (2) a multi-dimensional dataset profiling method that extracts and aggregates structural, semantic, and syntactic characteristics from prompt datasets, enabling comprehensive analysis of prompt distributions and patterns; (3) a controlled experimental framework for prompt sensitivity analysis by quantifying the impact of semantic reordering and delimiter modifications on LLM performance. Our experimental results validate the effectiveness of our taxonomy across these applications, demonstrating that PromptPrism provides a foundation for refining, profiling, and analyzing prompts.
Diffusion Language Model Inference with Monte Carlo Tree Search
Zheng Huang | Kiran Ramnath | Yueyan Chen | Aosong Feng | Sangmin Woo | Balasubramaniam Srinivasan | Zhichao Xu | Kang Zhou | Shuai Wang | Haibo Ding | Lin Lee Cheong
Findings of the Association for Computational Linguistics: EACL 2026
Zheng Huang | Kiran Ramnath | Yueyan Chen | Aosong Feng | Sangmin Woo | Balasubramaniam Srinivasan | Zhichao Xu | Kang Zhou | Shuai Wang | Haibo Ding | Lin Lee Cheong
Findings of the Association for Computational Linguistics: EACL 2026
Diffusion language models (DLMs) have recently emerged as a compelling alternative to autoregressive generation, offering parallel generation and improved global coherence. During inference, DLMs generate text by iteratively denoising masked sequences in parallel; however, determining which positions to unmask and which tokens to commit forms a large combinatorial search problem. Existing inference methods approximate this search using heuristics, which often yield suboptimal decoding paths; other approaches instead rely on additional training to guide token selection. To introduce a principled search mechanism for DLMs inference, we introduce MEDAL, an inference-time scaling framework that integrates Monte Carlo Tree SEarch initialization for Diffusion LAnguage Model inference. We employ Monte Carlo Tree Search at the initialization stage to explore promising unmasking trajectories, providing a robust starting point for subsequent refinement. This design enables efficient inference-time scaling, allowing generation quality to improve as the search budget increases, without additional training. Across multiple benchmarks, MEDAL achieves up to 22.0% improvement over existing inference strategies, establishing a new paradigm for search-based inference in DLMs.
BoundRL: Efficient Token-level Structured Text Segmentation through Reinforced Boundary Generation
Haoyuan Li | Zhengyuan Shen | Sullam Jeoung | Yueyan Chen | Jiayu Li | Qi Zhu | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Findings of the Association for Computational Linguistics: ACL 2026
Haoyuan Li | Zhengyuan Shen | Sullam Jeoung | Yueyan Chen | Jiayu Li | Qi Zhu | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Findings of the Association for Computational Linguistics: ACL 2026
Structured texts – from technical reports to AI prompts – increasingly require segmentation into semantically meaningful components. Such texts often contain elements beyond plain language, such as code snippets, which conventional sentence-level segmentation methods cannot handle effectively. To address this, we propose BoundRL, a novel approach that jointly performs efficient token-level text segmentation and label prediction for long structured texts. Instead of generating full texts for each segment, it generates only starting tokens and reconstructs the complete texts by locating these tokens within the original texts, thereby reducing inference costs by 90% and minimizing hallucination. To train the models for the boundary generation, BoundRL performs reinforcement learning with verifiable rewards (RLVR) that jointly optimizes document reconstruction fidelity and semantic alignment. It further mitigates entropy collapse by constructing intermediate candidates by perturbing segment boundaries and labels to create stepping stones toward higher-quality solutions. Experiments show that BoundRL enables small language models (1.7B parameters) to outperform few-shot prompting with much larger models as well as SFT and standard RLVR baselines on complex prompts used for LLM applications.
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
Yuqing Yang | Qi Zhu | Zhen Han | Boran Han | Zhengyuan Shen | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yuqing Yang | Qi Zhu | Zhen Han | Boran Han | Zhengyuan Shen | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e., incorrectly citing or omitting table values, despite understanding the table structure. Beyond final-answer accuracy, DREs directly compromise the correctness and reliability of intermediate reasoning steps. Yet prior studies have only offered limited, small-scale analyses. In this work, we present the first systematic evaluation of tabular data referencing errors across different models and tasks. Our results show that DREs occur across all tested models (1.7B to 20B parameters). Furthermore, we demonstrate that incorporating data referencing as a critic significantly improves answer accuracy up to 12.0%, through critic-based filtering and rejection sampling. Finally, we trained a lightweight 4B-parameter critic model that achieves an average F1 score of 78.2% in detecting both in-distribution and out-of-distribution DREs, and effectively assists inference for larger models.
SQL-Trail: Multi-Turn Reinforcement Learning with Interleaved Feedback for Text-to-SQL
Harper Hua | Zhen Han | Zhengyuan Shen | Meng-Chieh Lee | Sheng Guan | Qi Zhu | Sullam Jeoung | Yueyan Chen | Yunfei Bai | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Harper Hua | Zhen Han | Zhengyuan Shen | Meng-Chieh Lee | Sheng Guan | Qi Zhu | Sullam Jeoung | Yueyan Chen | Yunfei Bai | Shuai Wang | Vassilis N. Ioannidis | Huzefa Rangwala
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
While large language models (LLMs) have substantially improved Text-to-SQL generation, a pronounced gap remains between AI systems and human experts on challenging benchmarks such as BIRD-SQL. We argue this gap stems largely from the prevailing single-pass paradigm, which lacks the iterative reasoning, schema exploration, and error-correction behaviors that humans naturally employ. To address this limitation, we introduce SQL-Trail, a multi-turn reinforcement learning (RL) agentic framework for Text-to-SQL. Rather than producing a query in one shot, SQL-Trail interacts with the database environment and uses execution feedback to iteratively refine its predictions. Our approach centers on two key ideas: (i) an adaptive turn-budget allocation mechanism that scales the agent’s interaction depth to match question difficulty, and (ii) a composite reward panel that jointly incentivizes SQL correctness and efficient exploration. Across benchmarks, SQL-Trail sets a new state of the art and delivers strong data efficiency—up to 18× higher than prior single-pass RL state-of-the-art methods. Notably, our 7B and 14B models outperform substantially larger proprietary systems by 5% on average, underscoring the effectiveness of interactive, agentic workflows for robust Text-to-SQL generation.
2025
Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
Sangmin Woo | Kang Zhou | Yun Zhou | Shuai Wang | Sheng Guan | Haibo Ding | Lin Lee Cheong
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers)
Sangmin Woo | Kang Zhou | Yun Zhou | Shuai Wang | Sheng Guan | Haibo Ding | Lin Lee Cheong
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers)
Large Vision Language Models (LVLMs) often suffer from object hallucination, which undermines their reliability. Surprisingly, we find that simple object-based visual prompting—overlaying visual cues (e.g., bounding box, circle) on images—can significantly mitigate such hallucination; however, different visual prompts (VPs) vary in effectiveness. To address this, we propose Black-Box Visual Prompt Engineering (BBVPE), a framework to identify optimal VPs that enhance LVLM responses without needing access to model internals. Our approach employs a pool of candidate VPs and trains a router model to dynamically select the most effective VP for a given input image. This black-box approach is model-agnostic, making it applicable to both open-source and proprietary LVLMs. Evaluations on benchmarks such as POPE and CHAIR demonstrate that BBVPE effectively reduces object hallucination.
Aligning to Constraints for Data-Efficient Language Model Customization
Fei Wang | Chao Shang | Shuai Wang | Sarthak Jain | Qiang Ning | Bonan Min | Vittorio Castelli | Yassine Benajiba | Dan Roth
Findings of the Association for Computational Linguistics: NAACL 2025
Fei Wang | Chao Shang | Shuai Wang | Sarthak Jain | Qiang Ning | Bonan Min | Vittorio Castelli | Yassine Benajiba | Dan Roth
Findings of the Association for Computational Linguistics: NAACL 2025
General-purpose language models (LMs) are aligned to diverse user intents, but fall short when it comes to specific applications. While finetuning is the default method for customized alignment, human annotations are often unavailable in various customization scenarios. Based on the observation that one of the main issues of LM customization is constraint adherence, we investigate the feasibility of using constraints as a bridge from general LMs to customized ones. We investigate common constraints in NLP tasks, categorize them into three classes based on the types of their arguments, and propose a unified framework, ACT (Aligning to ConsTraints), to automatically produce supervision signals for user alignment with constraints. Specifically, ACT uses constraint verifiers, which are typically easy to implement in practice, to compute constraint satisfaction rate (CSR) of each response. It samples multiple responses for each prompt and collect preference labels based on their CSR automatically. Subsequently, ACT adapts the LM to the target task through a ranking-based learning process. Experiments on fine-grained entity typing, abstractive summarization, and temporal question answering show that ACT is able to enhance LMs’ capability to adhere to different classes of constraints, thereby improving task performance comparable to or approaching that of finetuning with labeled data.
A Systematic Survey of Automatic Prompt Optimization Techniques
Kiran Ramnath | Kang Zhou | Sheng Guan | Soumya Smruti Mishra | Xuan Qi | Zhengyuan Shen | Shuai Wang | Sangmin Woo | Sullam Jeoung | Yawei Wang | Haozhu Wang | Han Ding | Yuzhe Lu | Zhichao Xu | Yun Zhou | Balasubramaniam Srinivasan | Qiaojing Yan | Yueyan Chen | Haibo Ding | Panpan Xu | Lin Lee Cheong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Kiran Ramnath | Kang Zhou | Sheng Guan | Soumya Smruti Mishra | Xuan Qi | Zhengyuan Shen | Shuai Wang | Sangmin Woo | Sullam Jeoung | Yawei Wang | Haozhu Wang | Han Ding | Yuzhe Lu | Zhichao Xu | Yun Zhou | Balasubramaniam Srinivasan | Qiaojing Yan | Yueyan Chen | Haibo Ding | Panpan Xu | Lin Lee Cheong
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Since the advent of large language models (LLMs), prompt engineering has been a crucial step for eliciting desired responses for various Natural Language Processing (NLP) tasks. However, prompt engineering remains an impediment for end users due to rapid advances in models, tasks, and associated best practices. To mitigate this, Automatic Prompt Optimization (APO) techniques have recently emerged that use various automated techniques to help improve the performance of LLMs on various tasks. In this paper, we present a comprehensive survey summarizing the current progress and remaining challenges in this field. We provide a formal definition of APO, a 5-part unifying framework, and then proceed to rigorously categorize all relevant works based on their salient features therein. We hope to spur further research guided by our framework.
Search
Fix author
Co-authors
- Yueyan Chen 5
- Lin Lee Cheong 4
- Haibo Ding 4
- Sullam Jeoung 4
- Zhengyuan Shen 4
- Sheng Guan 3
- Vassilis N. Ioannidis 3
- Huzefa Rangwala 3
- Sangmin Woo 3
- Kang Zhou 3
- Qi Zhu 3
- Zhen Han 2
- Kiran Ramnath 2
- Balasubramaniam Srinivasan 2
- Zhichao Xu 2
- Yun Zhou 2
- Yunfei Bai 1
- Yassine Benajiba 1
- Vittorio Castelli 1
- Han Ding 1
- Aosong Feng 1
- Boran Han 1
- Harper Hua 1
- Zheng Huang 1
- Sarthak Jain 1
- Meng-Chieh Lee 1
- Haoyuan Li 1
- Jiayu Li 1
- Yuzhe Lu 1
- Bonan Min 1
- Soumya Smruti Mishra 1
- Qiang Ning 1
- Xuan Qi 1
- Dan Roth 1
- Chao Shang 1
- Fei Wang 1
- Haozhu Wang 1
- Yawei Wang 1
- Panpan Xu 1
- Qiaojing Yan 1
- Yuqing Yang 1
- Yi Zhang 1