Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications

Oren Sultan, Alexander Khasin, Guy Shiran, Asnat Greenstein-Messica, Dafna Shahaf


Abstract
We present a practical distillation approach to fine-tune LLMs for invoking tools in real-time applications. We focus on visual editing tasks; specifically, we modify images and videos by interpreting user stylistic requests, specified in natural language (“golden hour”), using an LLM to select the appropriate tools and their parameters to achieve the desired visual effect.We found that proprietary LLMs such as GPT-3.5-Turbo show potential in this task, but their high cost and latency make them unsuitable for real-time applications.In our approach, we fine-tune a (smaller) student LLM with guidance from a (larger) teacher LLM and behavioral signals.We introduce offline metrics to evaluate student LLMs. Both online and offline experiments show that our student models manage to match the performance of our teacher model (GPT-3.5-Turbo), significantly reducing costs and latency.Lastly, we show that fine-tuning was improved by 25% in low-data regimes using augmentation.
Anthology ID:
2024.emnlp-industry.96
Volume:
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track
Month:
November
Year:
2024
Address:
Miami, Florida, US
Editors:
Franck Dernoncourt, Daniel Preoţiuc-Pietro, Anastasia Shimorina
Venue:
EMNLP
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
1286–1304
Language:
URL:
https://aclanthology.org/2024.emnlp-industry.96
DOI:
Bibkey:
Cite (ACL):
Oren Sultan, Alexander Khasin, Guy Shiran, Asnat Greenstein-Messica, and Dafna Shahaf. 2024. Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 1286–1304, Miami, Florida, US. Association for Computational Linguistics.
Cite (Informal):
Visual Editing with LLM-based Tool Chaining: An Efficient Distillation Approach for Real-Time Applications (Sultan et al., EMNLP 2024)
Copy Citation:
PDF:
https://aclanthology.org/2024.emnlp-industry.96.pdf
Poster:
 2024.emnlp-industry.96.poster.pdf
Presentation:
 2024.emnlp-industry.96.presentation.pdf
Video:
 2024.emnlp-industry.96.video.mp4