Yuheng Lu
Author directoryPapers on this page may belong to the following people: Yuheng Lu, Yuheng Lu
2026
Code-Switching Information Retrieval: Benchmarks, Analysis, and the Limits of Current Retrievers
Qingcheng Zeng | Yuheng Lu | Zeqi Zhou | Heli Qi | Puxuan Yu | Fuheng Zhao | Hitomi Yanaka | Weihao Xuan | Naoto Yokoya
Findings of the Association for Computational Linguistics: ACL 2026
Qingcheng Zeng | Yuheng Lu | Zeqi Zhou | Heli Qi | Puxuan Yu | Fuheng Zhao | Hitomi Yanaka | Weihao Xuan | Naoto Yokoya
Findings of the Association for Computational Linguistics: ACL 2026
Code-switching is a pervasive linguistic phenomenon in global communication, yet modern information retrieval systems remain predominantly designed for, and evaluated within, monolingual contexts. To bridge this critical disconnect, we present a holistic study dedicated to code-switching IR. We introduce CSR-L (Code-Switching Retrieval benchmark-Lite), constructing a dataset via human annotation to capture the authentic naturalness of mixed-language queries. Our evaluation across statistical, dense, and late-interaction paradigms reveals that code-switching acts as a fundamental performance bottleneck, degrading the effectiveness of even robust multilingual models. We demonstrate that this failure stems from substantial divergence in the embedding space between pure and code-switched text. Scaling this investigation, we propose CS-MTEB, a comprehensive benchmark covering 11 diverse tasks, where we observe performance declines of up to 27%. Finally, we show that standard multilingual techniques like vocabulary expansion are insufficient to resolve these deficits completely. These findings underscore the fragility of current systems and establish code-switching as a crucial frontier for future IR optimization.
2025
Controlled Low-Rank Adaptation with Subspace Regularization for Continued Training on Large Language Models
Yuheng Lu | Bingshuo Qian | Caixia Yuan | Huixing Jiang | Xiaojie Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Yuheng Lu | Bingshuo Qian | Caixia Yuan | Huixing Jiang | Xiaojie Wang
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Large language models (LLMs) exhibit remarkable capabilities in natural language processing but face catastrophic forgetting when learning new tasks, where adaptation to a new domain leads to a substantial decline in performance on previous tasks. In this paper, we propose Controlled LoRA (CLoRA), a subspace regularization method on LoRA structure. Aiming to reduce the scale of output change while introducing minimal constraint on model capacity, CLoRA imposes constraints on the direction of updating matrix’s null space. Experimental results on one-stage LLM finetuning tasks and continual learning settings highlight the superiority of CLoRA as an effective parameter-efficient finetuning method with catastrophic forgetting mitigating. Further investigation for model parameters indicates that CLoRA effectively balances the trade-off between model capacity and degree of forgetting. The code for implementing CLoRA will be publicly available.