Vertical Routing: A Cost-Efficient Collaboration Routing Framework

Si Shen, Peijun Shen, Danhao Zhu


Abstract
Routing Large and Small Language Models (LLMs SLMs) is commonly framed as query-level difficulty prediction, yet lightweight routers are often unreliable and stronger evaluators introduce non-trivial overhead. We propose VERTICAL ROUTING, a stage-level collaboration framework that avoids monolithic difficulty prediction by allocating the large model to critical subtasks and delegating the remaining generation to a small model. VERTICAL ROUTING instantiates two templates: (i) domain-specific templates that decompose a task into ordered stages with criticality scores, and (ii) a robust default template that follows a prefix-first prior for general queries. Under a token budget, we allocate large-model capacity to the most critical stages, or to a budgeted prefix under the default template. Experiments show that VERTICAL ROUTING outperforms the strongest baseline (RouterDC) by 2.0 points in the averaged metric, while significantly reducing token usage by 48.9% and large-model output share by 41.7%. Additionally, it enhances stability by lowering the standard deviation by 87.5%. These results highlight its advantages in accuracy, efficiency, and robustness.1
Anthology ID:
2026.tacl-1.104
Volume:
Transactions of the Association for Computational Linguistics, Volume 14
Month:
Year:
2026
Address:
Cambridge, MA
Venue:
TACL
SIG:
Publisher:
MIT Press
Note:
Pages:
2301–2317
Language:
URL:
https://aclanthology.org/2026.tacl-1.104/
DOI:
10.1162/tacl.a.802
Bibkey:
Cite (ACL):
Si Shen, Peijun Shen, and Danhao Zhu. 2026. Vertical Routing: A Cost-Efficient Collaboration Routing Framework. Transactions of the Association for Computational Linguistics, 14:2301–2317.
Cite (Informal):
Vertical Routing: A Cost-Efficient Collaboration Routing Framework (Shen et al., TACL 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.tacl-1.104.pdf