TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use

Junjie Ye (叶俊杰); Yilong Wu; Sixian Li; Yuming Yang; Zhiheng Xi; Tao Gui; Qi Zhang; Xuan-Jing Huang (黄萱菁); Peng Wang; Zhongchao Shi; Jianping Fan; Zhengyin Du

doi:10.18653/v1/2025.findings-emnlp.15

TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use

Junjie Ye, Yilong Wu, Sixian Li, Yuming Yang, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan, Zhengyin Du

Abstract

Large language models (LLMs) achieve remarkable advancements by leveraging tools to interact with environments, a critical step toward generalized AI. However, the standard supervised fine-tuning (SFT) approach, which relies on large-scale datasets, often overlooks task-specific characteristics in tool use, leading to performance bottlenecks. To address this issue, we analyze three existing LLMs and uncover key insights: training data can inadvertently impede tool-use behavior, token importance is distributed unevenly, and errors in tool calls fall into a small set of categories. Building on these findings, we propose TL-Training, a task-feature-based framework that mitigates the effects of suboptimal training data, dynamically adjusts token weights to prioritize key tokens during SFT, and incorporates a robust reward mechanism tailored to error categories, optimized through proximal policy optimization. We validate TL-Training by training CodeLLaMA-2-7B and evaluating it on four open-source test sets. Our results demonstrate that the LLM trained by our method matches or surpasses both open- and closed-source LLMs in tool-use performance using only 1,217 training data points. Additionally, our method enhances robustness in noisy environments and improves general task performance, offering a scalable and efficient paradigm for tool-use training in LLMs. Code and data are available at https://github.com/Junjie-Ye/TL-Training.

Anthology ID:: 2025.findings-emnlp.15
Volume:: Findings of the Association for Computational Linguistics: EMNLP 2025
Month:: November
Year:: 2025
Address:: Suzhou, China
Editors:: Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, Violet Peng
Venue:: Findings
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 239–258
Language:
URL:: https://aclanthology.org/2025.findings-emnlp.15/
DOI:: 10.18653/v1/2025.findings-emnlp.15
Bibkey:
Cite (ACL):: Junjie Ye, Yilong Wu, Sixian Li, Yuming Yang, Zhiheng Xi, Tao Gui, Qi Zhang, Xuanjing Huang, Peng Wang, Zhongchao Shi, Jianping Fan, and Zhengyin Du. 2025. TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 239–258, Suzhou, China. Association for Computational Linguistics.
Cite (Informal):: TL-Training: A Task-Feature-Based Framework for Training Large Language Models in Tool Use (Ye et al., Findings 2025)
Copy Citation:
PDF:: https://aclanthology.org/2025.findings-emnlp.15.pdf
Checklist:: 2025.findings-emnlp.15.checklist.pdf

PDF Cite Search Checklist Fix data