Liwen Zhang
Author directoryOther people with similar names: Liwen Zhang
Unverified author pages with similar names: Liwen Zhang
2026
BrowseConf: Confidence-Guided Test-Time Scaling for Web Agents
Litu Ou | Kuan Li | Huifeng Yin | Liwen Zhang | Zhongwang Zhang | Xixi Wu | Rui Ye | Zile Qiao | Yong Jiang | Pengjun Xie | Fei Huang | Jingren Zhou
Findings of the Association for Computational Linguistics: ACL 2026
Litu Ou | Kuan Li | Huifeng Yin | Liwen Zhang | Zhongwang Zhang | Xixi Wu | Rui Ye | Zile Qiao | Yong Jiang | Pengjun Xie | Fei Huang | Jingren Zhou
Findings of the Association for Computational Linguistics: ACL 2026
Confidence in LLMs is a useful indicator of model uncertainty and answer reliability. Existing work mainly focused on single-turn scenarios, while research on confidence in complex multi-turn interactions is limited. In this paper, we investigate whether LLM-based search agents have the ability to communicate their own confidence through verbalized confidence scores after long sequences of actions, a significantly more challenging task compared to outputting confidence in a single interaction. Experimenting on open-source agentic models, we first find that models exhibit much higher task accuracy at high confidence while having near-zero accuracy when confidence is low. Based on this observation, we propose Test-Time Scaling (TTS) methods that use confidence scores to determine answer quality, encourage the model to try again until reaching a satisfactory confidence level. Results show that our proposed methods significantly reduce token consumption while demonstrating competitive performance compared to baseline fixed budget TTS methods.
WebAnchor: Anchoring Agent Planning to Stabilize Long-Horizon Web Reasoning
Xinmiao Yu | Liwen Zhang | Xiaocheng Feng | Pengjun Xie | Jingren Zhou | Bing Qin | Yong Jiang
Findings of the Association for Computational Linguistics: ACL 2026
Xinmiao Yu | Liwen Zhang | Xiaocheng Feng | Pengjun Xie | Jingren Zhou | Bing Qin | Yong Jiang
Findings of the Association for Computational Linguistics: ACL 2026
Large Language Model(LLM)-based agents have shown strong capabilities in web information seeking, with reinforcement learning (RL) becoming a key optimization paradigm. However, planning remains a bottleneck, as existing methods struggle with long-horizon strategies. Our analysis reveals a critical phenomenon—plan anchor—where the first reasoning step disproportionately impacts downstream behavior in long-horizon web reasoning tasks. Current RL algorithms, fail to account for this by uniformly distributing rewards across the trajectory.To address this, we propose Anchor-GRPO, a two-stage RL framework that decouples planning and execution. In Stage 1, the agent optimizes its first-step planning using fine-grained rubrics derived from self-play experiences and human calibration. In Stage 2, execution is aligned with the initial plan through sparse rewards, ensuring stable and efficient tool usage. We evaluate Anchor-GRPO on four benchmarks: BrowseComp, BrowseComp-Zh, GAIA, and XBench-DeepSearch. Across models from 3B to 30B, Anchor-GRPO outperforms baseline GRPO and First-step GRPO, improving task success and tool efficiency. Notably, WebAnchor-30B achieves 46.0% pass@1 on BrowseComp and 76.4% on GAIA. Anchor-GRPO also demonstrates strong scalability, getting higher accuracy as model size and context length increase.
Nested Browser-Use Learning for Agentic Information Seeking
Baixuan Li | Jialong Wu | Wenbiao Yin | Kuan Li | Zhongwang Zhang | Huifeng Yin | Zhengwei Tao | Liwen Zhang | Pengjun Xie | Jingren Zhou | Yong Jiang | Wentao Zhang | Zhiqiang Gao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Baixuan Li | Jialong Wu | Wenbiao Yin | Kuan Li | Zhongwang Zhang | Huifeng Yin | Zhengwei Tao | Liwen Zhang | Pengjun Xie | Jingren Zhou | Yong Jiang | Wentao Zhang | Zhiqiang Gao
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Information-seeking (IS) agents have achieved strong performance across a range of wide and deep search tasks, yet their tool use remains largely restricted to API-level snippet retrieval and URL-based page fetching, limiting access to the richer information available through real browsing. While full browser interaction could unlock deeper capabilities, its fine-grained control and verbose page content returns introduce substantial complexity for ReAct-style function-calling agents. To bridge this gap, we propose Nested Browser-Use Learning (NestBrowse), which introduces a minimal and complete browser-action framework that decouples interaction control from page exploration through a nested structure. This design simplifies agentic reasoning while enabling effective deep-web information acquisition. Empirical results on challenging deep IS benchmarks demonstrate that NestBrowse offers clear benefits in practice. Further in-depth analyses underscore its efficiency.
2025
EvolveSearch: An Iterative Self-Evolving Search Agent
Ding-Chu Zhang | Yida Zhao | Jialong Wu | Liwen Zhang | Baixuan Li | Wenbiao Yin | Yong Jiang | Yu-Feng Li | Kewei Tu | Pengjun Xie | Fei Huang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Ding-Chu Zhang | Yida Zhao | Jialong Wu | Liwen Zhang | Baixuan Li | Wenbiao Yin | Yong Jiang | Yu-Feng Li | Kewei Tu | Pengjun Xie | Fei Huang
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
The rapid advancement of large language models (LLMs) has transformed the landscape of agentic information seeking capabilities through the integration of tools such as search engines and web browsers. However, current mainstream approaches for enabling LLM web search proficiency face significant challenges: supervised fine-tuning struggles with data production in open-search domains, while RL converges quickly, limiting their data utilization efficiency. To address these issues, we propose EvolveSearch, a novel iterative self-evolution framework that combines SFT and RL to enhance agentic web search capabilities without any external human-annotated reasoning data. Extensive experiments on seven multi-hop question-answering (MHQA) benchmarks demonstrate that EvolveSearch consistently improves performance across iterations, ultimately achieving an average improvement of 4.7% over the current state-of-the-art across seven benchmarks, opening the door to self-evolution agentic capabilities in open web search domains.