Yu Su
Author directoryOther people with similar names: Yu Su
Unverified author pages with similar names: Yu Su
2026
OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulation
Ziyi Wang | Yuxuan Lu | Wenbo Li | Amirali Amini | Bo Sun | Yakov Bart | Weimin Lyu | Jiri Gesi | Tian Wang | Jing Huang | Yu Su | Upol Ehsan | Malihe Alikhani | Toby Jia-Jun Li | Lydia Chilton | Dakuo Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Ziyi Wang | Yuxuan Lu | Wenbo Li | Amirali Amini | Bo Sun | Yakov Bart | Weimin Lyu | Jiri Gesi | Tian Wang | Jing Huang | Yu Su | Upol Ehsan | Malihe Alikhani | Toby Jia-Jun Li | Lydia Chilton | Dakuo Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Can Large Language models (LLMs) accurately simulate the next web action of a specific user? While LLMs have shown promising capabilities in generating believable human behaviors, evaluating their ability to mimic real user behaviors remains an open challenge, largely due to the lack of high-quality, publicly available datasets that capture both the observable actions and the internal reasoning of an actual human user. To address this gap, we introduce OPeRA, a novel dataset of Observation, Persona, Rationale, and Action collected from real human participants during online shopping sessions. OPeRA is the first public dataset that comprehensively captures: user personas, browser observations, fine-grained web actions, and self-reported just-in-time rationales. We developed both an online questionnaire and a custom browser plugin to gather this dataset with high fidelity. Using OPeRA, we establish the first benchmark to evaluate how well current LLMs can predict a specific user’s next action and rationale with a given persona and <observation, action, rationale> history. This dataset lays the groundwork for future research into LLM agents that aim to act as personalized digital twins for human.
2025
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
Nishant Subramani | Jason Eisner | Justin Svegliato | Benjamin Van Durme | Yu Su | Sam Thomson
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Nishant Subramani | Jason Eisner | Justin Svegliato | Benjamin Van Durme | Yu Su | Sam Thomson
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
Tool-using agents that act in the world need to be both useful and safe. Well-calibrated model confidences can be used to weigh the risk versus reward of potential actions, but prior work shows that many models are poorly calibrated. Inspired by interpretability literature exploring the internals of models, we propose a novel class of model-internal confidence estimators (MICE) to better assess confidence when calling tools. MICE first decodes from each intermediate layer of the language model using logit lens and then computes similarity scores between each layer’s generation and the final output. These features are fed into a learned probabilistic classifier to assess confidence in the decoded output. On the simulated trial and error (STE) tool-calling dataset using Llama3 models, we find that MICE beats or matches the baselines on smoothed expected calibration error. Using MICE confidences to determine whether to call a tool significantly improves over strong baselines on a new metric, expected tool-calling utility. Further experiments show that MICE is sample-efficient, can generalize zero-shot to unseen APIs, and results in higher tool-calling utility in scenarios with varying risk levels. Our code is open source, available at https://github.com/microsoft/mice_for_cats.
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents
Vardaan Pahuja | Yadong Lu | Corby Rosset | Boyu Gou | Arindam Mitra | Spencer Whitehead | Yu Su | Ahmed Hassan Awadallah
Findings of the Association for Computational Linguistics: ACL 2025
Vardaan Pahuja | Yadong Lu | Corby Rosset | Boyu Gou | Arindam Mitra | Spencer Whitehead | Yu Su | Ahmed Hassan Awadallah
Findings of the Association for Computational Linguistics: ACL 2025
Recent success in large multimodal models (LMMs) has sparked promising applications of agents capable of autonomously completing complex web tasks. While open-source LMM agents have made significant advances in offline evaluation benchmarks, their performance still falls substantially short of human-level capabilities in more realistic online settings. A key bottleneck is the lack of diverse and large-scale trajectory-level datasets across various domains, which are expensive to collect. In this paper, we address this challenge by developing a scalable recipe to synthesize the largest and most diverse trajectory-level dataset to date, containing over 94K successful multimodal web trajectories, spanning 49K unique URLs, 720K screenshots, and 33M web elements. In particular, we leverage extensive web exploration and refinement to obtain diverse task intents. The average cost is 28 cents per successful trajectory, making it affordable to a wide range of users in the community. Leveraging this dataset, we train Explorer, a multimodal web agent, and demonstrate strong performance on both offline and online web agent benchmarks such as Mind2Web-Live, Multimodal-Mind2Web, and MiniWob++. Additionally, our experiments highlight data scaling as a key driver for improving web agent capabilities. We hope this study makes state-of-the-art LMM-based agent research at a larger scale more accessible.
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Xiang Yue | Tianyu Zheng | Yuansheng Ni | Yubo Wang | Kai Zhang | Shengbang Tong | Yuxuan Sun | Botao Yu | Ge Zhang | Huan Sun | Yu Su | Wenhu Chen | Graham Neubig
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Xiang Yue | Tianyu Zheng | Yuansheng Ni | Yubo Wang | Kai Zhang | Shengbang Tong | Yuxuan Sun | Botao Yu | Ge Zhang | Huan Sun | Yu Su | Wenhu Chen | Graham Neubig
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models’ true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly “see” and “read” simultaneously, testing a core human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future multimodal research.
Completing A Systematic Review in Hours instead of Months with Interactive AI Agents
Rui Qiu | Shijie Chen | Yu Su | Po-Yin Yen | Han Wei Shen
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Rui Qiu | Shijie Chen | Yu Su | Po-Yin Yen | Han Wei Shen
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Systematic reviews (SRs) are vital for evidence-based practice in high stakes disciplines, such as healthcare, but are often impeded by intensive labors and lengthy processes that can take months to complete. Due to the high demand for domain expertise, existing automatic summarization methods fail to accurately identify relevant studies and generate high-quality summaries. To that end, we introduce InsightAgent, a human-centered interactive AI agent powered by large language models that revolutionize this workflow. InsightAgent partitions a large literature corpus based on semantics and employs a multi-agent design for more focused processing of literature, leading to significant improvement in the quality of generated SRs. InsightAgent also provides intuitive visualizations of the corpus and agent trajectories, allowing users to effortlessly monitor the actions of the agent and provide real-time feedback based on their expertise. Our user studies with 9 medical professionals demonstrate that the visualization and interaction mechanisms can effectively improve the quality of synthesized SRs by 27.2%, reaching 79.7% of human-written quality. At the same time, user satisfaction is improved by 34.4%. With InsightAgent, it only takes a clinician about 1.5 hours, rather than months, to complete a high-quality systematic review.
2024
Middleware for LLMs: Tools Are Instrumental for Language Agents in Complex Environments
Yu Gu | Yiheng Shu | Hao Yu | Xiao Liu | Yuxiao Dong | Jie Tang | Jayanth Srinivasa | Hugo Latapie | Yu Su
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Yu Gu | Yiheng Shu | Hao Yu | Xiao Liu | Yuxiao Dong | Jie Tang | Jayanth Srinivasa | Hugo Latapie | Yu Su
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
The applications of large language models (LLMs) have expanded well beyond the confines of text processing, signaling a new era where LLMs are envisioned as generalist agents capable of operating within complex environments. These environments are often highly expansive, making it impossible for the LLM to process them within its short-term memory. Motivated by recent research on extending the capabilities of LLMs with tools, we seek to investigate the intriguing potential of tools to augment LLMs in handling such complexity by introducing a novel class of tools, termed middleware, to aid in the proactive exploration within these massive environments. Such specialized tools can serve as a middleware layer shielding the LLM from environmental complexity. In two representative complex environments—knowledge bases (KBs) and databases—we demonstrate the significant potential of augmenting language agents with tools in complex environments. Notably, equipped with the middleware, GPT-4 achieves 2.8X the performance of the best baseline in tasks requiring access to database content and 2.2X in KB tasks. Our findings illuminate the path for advancing language agents in real-world applications.
WebOlympus: An Open Platform for Web Agents on Live Websites
Boyuan Zheng | Boyu Gou | Scott Salisbury | Zheng Du | Huan Sun | Yu Su
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
Boyuan Zheng | Boyu Gou | Scott Salisbury | Zheng Du | Huan Sun | Yu Su
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
Web agents are emerging as powerful tools capable of performing complex tasks across diverse web environments. The rapid development of large multimodal models is further enhancing this advancement. However, there is a lack of standardized and user-friendly tools for research and development, as well as experimental platforms on live websites. To address this challenge, we present WebOlympus, an open platform for web agents operating on live websites. WebOlympus offers a Chrome extension-based UI, enabling users without programming experience to easily utilize the platform. It allows users to run web agents with various designs using only a few lines of code or simple clicks on the Chrome extension. To ensure the trustworthiness of web agents, a safety monitor module that prevents harmful actions through human supervision or model-based control is incorporated. WebOlympus supports diverse applications, including annotation interfaces for web agent trajectories and data crawling.
2022
Bridging the Generalization Gap in Text-to-SQL Parsing with Schema Expansion
Chen Zhao | Yu Su | Adam Pauls | Emmanouil Antonios Platanios
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Chen Zhao | Yu Su | Adam Pauls | Emmanouil Antonios Platanios
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Text-to-SQL parsers map natural language questions to programs that are executable over tables to generate answers, and are typically evaluated on large-scale datasets like Spider (Yu et al., 2018). We argue that existing benchmarks fail to capture a certain out-of-domain generalization problem that is of significant practical importance: matching domain specific phrases to composite operation over columns. To study this problem, we first propose a synthetic dataset along with a re-purposed train/test split of the Squall dataset (Shi et al., 2020) as new benchmarks to quantify domain generalization over column operations, and find existing state-of-the-art parsers struggle in these benchmarks. We propose to address this problem by incorporating prior domain knowledge by preprocessing table schemas, and design a method that consists of two components: schema expansion and schema pruning. This method can be easily applied to multiple existing base parsers, and we show that it significantly outperforms baseline parsers on this domain generalization problem, boosting the underlying parsers’ overall performance by up to 13.8% relative accuracy gain (5.1% absolute) on the new Squall data split.
Search
Fix author
Co-authors
- Boyu Gou 2
- Huan Sun 2
- Malihe Alikhani 1
- Amirali Amini 1
- Yakov Bart 1
- Shijie Chen 1
- Wenhu Chen 1
- Lydia Chilton 1
- Yuxiao Dong 1
- Zheng Du 1
- Upol Ehsan 1
- Jason Eisner 1
- Jiri Gesi 1
- Yu Gu (谷峪) 1
- Ahmed Hassan 1
- Jing Huang 1
- Hugo Latapie 1
- Toby Jia-Jun Li 1
- Wenbo Li 1
- Xiao Liu 1
- Yadong Lu 1
- Yuxuan Lu 1
- Weimin Lyu 1
- Arindam Mitra 1
- Graham Neubig 1
- Yuansheng Ni 1
- Vardaan Pahuja 1
- Adam Pauls 1
- Emmanouil Antonios Platanios 1
- Rui Qiu 1
- Corby Rosset 1
- Scott Salisbury 1
- Han Wei Shen 1
- Yiheng Shu 1
- Jayanth Srinivasa 1
- Nishant Subramani 1
- Bo Sun 1
- Yuxuan Sun 1
- Justin Svegliato 1
- Jie Tang 1
- Sam Thomson 1
- Shengbang Tong 1
- Benjamin Van Durme 1
- Dakuo Wang 1
- Tian Wang 1
- Yubo Wang 1
- Ziyi Wang 1
- Spencer Whitehead 1
- Po-Yin Yen 1
- Botao Yu 1
- Hao Yu (余浩) 1
- Xiang Yue 1
- Ge Zhang 1
- Kai Zhang 1
- Chen Zhao 1
- Boyuan Zheng 1
- Tianyu Zheng 1