Zhihan Li
Author directory2026
LoRA Fine-Tuning of English–Norwegian NMT for the Oil & Gas Industry
Xiaojing Yang | Zhihan Li | Gege Sun | Mengyue Li | Meriem Beloucif
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Xiaojing Yang | Zhihan Li | Gege Sun | Mengyue Li | Meriem Beloucif
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Adapting large language models to specialized domains remains challenging due to the computational cost of full finetuning and the limited availability of domain-specific parallel data. We present a systematic framework for parameter-efficient domain adaptation using Low-Rank Adaptation (LoRA) geared towards efficient learning in low-resource scenarios. Our method combines data-scaling analysis, dual-track hyperparameter optimization, and competitive benchmarking. We evaluate our approach on the low-resource English–Norwegian petroleum translation domain using a distilled version of NLLB and parallel data from the Norwegian Petroleum Directorate. Our adapted model achieves 61.48 BLEU (+24.62 over the base model) and 0.9298 COMET, while updating <0.4% of parameters. Our results provide a reproducible and computationally efficient blueprint for domain adaptation in neural machine translation, particularly for specialized and resource-constrained domains.
2024
What Are the Implications of Your Question? Non-Information Seeking Question-Type Identification in CNN Transcripts
Yao Sun | Anastasiia Tatlubaeva | Zhihan Li | Chester Palen-Michel
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Yao Sun | Anastasiia Tatlubaeva | Zhihan Li | Chester Palen-Michel
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Non-information seeking questions (NISQ) capture the subtle dynamics of human discourse. In this work, we utilize a dataset of over 1,500 information-seeking question(ISQ) and NISQ to evaluate human and machine performance on classifying fine-grained NISQ types. We introduce the first publicly available corpus focused on annotating both ISQs and NISQs as an initial benchmark. Additionally, we establish competitive baselines by assessing diverse systems, including Generative Pre-Trained Transformer Language models, on a new question classification task. Our results demonstrate the inherent complexity of making nuanced NISQ distinctions. The dataset is publicly available at https://github.com/YaoSun0422/NISQ_dataset.git
2023
Model-Agnostic Meta-Learning for Natural Language Understanding Tasks in Finance
Bixing Yan | Shaoling Chen | Yuxuan He | Zhihan Li
Proceedings of the Fifth Workshop on Financial Technology and Natural Language Processing and the Second Multimodal AI For Financial Forecasting
Bixing Yan | Shaoling Chen | Yuxuan He | Zhihan Li
Proceedings of the Fifth Workshop on Financial Technology and Natural Language Processing and the Second Multimodal AI For Financial Forecasting