Jiahuan Li
Author directoryOther people with similar names: Jiahuan Li
Unverified author pages with similar names: Jiahuan Li
2026
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
Zhihao Xu | Rumei Li | Jiahuan Li | Rongxiang Weng | Jingang Wang | Xunliang Cai | Xiting Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Zhihao Xu | Rumei Li | Jiahuan Li | Rongxiang Weng | Jingang Wang | Xunliang Cai | Xiting Wang
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Enabling Large Language Models (LLMs) to effectively utilize tools in multi-turn interactions is essential for building capable autonomous agents. However, acquiring diverse and realistic multi-turn tool-use data remains a significant challenge. In this work, we propose a novel text-based paradigm. We observe that textual corpora naturally contain rich, multi-step problem-solving experiences, which can serve as an untapped, scalable, and authentic data source for multi-turn tool-use tasks. Based on this insight, we introduce GEM, a data synthesis pipeline that enables the generation and extraction of multi-turn tool-use trajectories from text corpora through a four-stage process: relevance filtering, workflow tool extraction, trajectory grounding, and complexity refinement. To reduce the computational cost, we further train a specialized Trajectory Synthesizer via supervised fine-tuning. This model distills the complex generation pipeline into an efficient, end-to-end trajectory generator. Experiments demonstrate that our GEM-32B achieve a 14.9% improvement on the BFCL V3 Multi-turn benchmark. Our models partially surpass the performance of models trained on -bench (Airline and Retail) in-domain data, highlighting the superior generalization capability derived from our text-based synthesis paradigm. Notably, our Trajectory Synthesizer matches the quality of the full pipeline while significantly reducing inference latency and costs.
2025
Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training
Zhijun Wang | Jiahuan Li | Hao Zhou | Rongxiang Weng | Jingang Wang | Xin Huang | Xue Han | Junlan Feng | Chao Deng | Shujian Huang
Findings of the Association for Computational Linguistics: ACL 2025
Zhijun Wang | Jiahuan Li | Hao Zhou | Rongxiang Weng | Jingang Wang | Xin Huang | Xue Han | Junlan Feng | Chao Deng | Shujian Huang
Findings of the Association for Computational Linguistics: ACL 2025
Large language models (LLMs) exhibit remarkable multilingual capabilities despite the extreme language imbalance in the pre-training data. In this paper, we closely examine the reasons behind this phenomenon, focusing on the pre-training corpus. We find that the existence of code-switching, alternating between different languages within a context, is key to multilingual capabilities. We conduct an analysis to investigate code-switching in the pre-training corpus, examining its presence and categorizing it into four types within two quadrants. We then assess its impact on multilingual performance. These types of code-switching data are unbalanced in proportions and demonstrate different effects on facilitating language transfer. To better explore the power of code-switching for language alignment during pre-training, we investigate the strategy of synthetic code-switching. We continuously scale up the synthetic code-switching data and observe remarkable improvements in both benchmarks and representation space. Extensive experiments indicate that incorporating synthetic code-switching data enables better language alignment and generalizes well to high, medium, and low-resource languages with pre-training corpora of varying qualities.