Dawei Zhu
Author directoryOther people with similar names: Dawei Zhu
Unverified author pages with similar names: Dawei Zhu
2026
What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary Translation
Shaomu Tan | Dawei Zhu | Ke Tran | Michael Denkowski | Sony Trenous | Leonardo F. R. Ribeiro | Bill Byrne | Felix Hieber
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Shaomu Tan | Dawei Zhu | Ke Tran | Michael Denkowski | Sony Trenous | Leonardo F. R. Ribeiro | Bill Byrne | Felix Hieber
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Iterative refinement is a simple inference-time strategy for machine translation: given an initial translation, an LLM revises it without additional training. Yet document-scale refinement remains poorly understood: 1) which pipelines work best, 2) what quality dimensions improve, and 3) how refiners behave. In this paper, we present a systematic study of document-level literary translation, covering six LLMs and seven language pairs. Across nine translation-refinement granularity combinations and five refinement strategies, a) we find a robust recipe: document-level MT followed by segment-level refinement yields the strongest and most stable improvements. In our setting, doc-level refinement often makes fewer edits and leads to smaller or less reliable gains. Surprisingly, a simple general refinement prompt consistently outperforms error-specific prompting and evaluate-then-refine schemes. b) Fine-grained MQM analyses and professional-translator evaluation show that gains come primarily from fluency, with limited improvements in adequacy. c) Probing translator-refiner strength interactions suggests refinement behaves less like targeted post-editing and more like projecting outputs toward the refiner’s learned distribution while remaining anchored to the initial translation.
2025
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation
Yirong Sun | Dawei Zhu | Yanjun Chen | Erjia Xiao | Xinghao Chen | Xiaoyu Shen
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop)
Yirong Sun | Dawei Zhu | Yanjun Chen | Erjia Xiao | Xinghao Chen | Xiaoyu Shen
Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop)
Large language models (LLMs) have excelled in various NLP tasks, including machine translation (MT), yet most studies focus on sentence-level translation. This work investigates the inherent capability of instruction-tuned LLMs for document-level translation (docMT). Unlike prior approaches that require specialized techniques, we evaluate LLMs by directly prompting them to translate entire documents in a single pass. Our results show that this method improves translation quality compared to translating sentences separately, even without document-level fine-tuning. However, this advantage is not reflected in BLEU scores, which often favor sentence-based translations. We propose using the LLM-as-a-judge paradigm for evaluation, where GPT-4 is used to assess document coherence, accuracy, and fluency in a more nuanced way than n-gram-based metrics. Overall, our work demonstrates that instruction-tuned LLMs can effectively leverage document context for translation. However, we caution against using BLEU scores for evaluating docMT, as they often provide misleading outcomes, failing to capture the quality of document-level translation.
From Calculation to Adjudication: Examining LLM Judges on Mathematical Reasoning Tasks
Andreas Stephan | Dawei Zhu | Matthias Aßenmacher | Xiaoyu Shen | Benjamin Roth
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
Andreas Stephan | Dawei Zhu | Matthias Aßenmacher | Xiaoyu Shen | Benjamin Roth
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlation with human judgments on generative tasks such as summarization or machine translation. In contrast, we study LLM judges on mathematical reasoning tasks. These tasks require multi-step reasoning, and the correctness of their solutions is verifiable, enabling a more objective evaluation. We perform a detailed performance analysis and find that easy samples are easy to judge, and difficult samples are difficult to judge. Our analysis uncovers a strong correlation between judgment performance and the candidate model task performance, indicating that judges tend to favor higher-quality models even if their answer is incorrect. As a consequence, we test whether we can predict the behavior of LLM judges using simple features such as part-of-speech tags and find that we can correctly predict 70%-75% of judgments. We conclude this study by analyzing practical use cases, showing that LLM judges consistently detect the on-average better model but largely fail if we use them to improve task performance.
Language models can learn implicit multi-hop reasoning, but only if they have lots of training data
Yuekun Yao | Yupei Du | Dawei Zhu | Michael Hahn | Alexander Koller
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Yuekun Yao | Yupei Du | Dawei Zhu | Michael Hahn | Alexander Koller
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Implicit reasoning is the ability of a language model to solve multi-hop reasoning tasks in a single forward pass, without chain of thought.We investigate this capability using GPT2-style language models trained from scratch on controlled k-hop reasoning datasets (k = 2, 3, 4). We show that while such models can indeed learn implicit k-hop reasoning,the required training data grows exponentially in k, and the requirednumber of transformer layers grows linearly in k.We offer a theoretical explanation for why this depth growth is necessary.We further find that the data requirement can be mitigated, but not eliminated,through curriculum learning.
Same evaluation, more tokens: On the effect of input length for machine translation evaluation using Large Language Models
Tobias Domhan | Dawei Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Tobias Domhan | Dawei Zhu
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Accurately evaluating machine-translated text remains a long-standing challenge, particularly for long documents. Recent work has shown that large language models (LLMs) can serve as reliable and interpretable sentence-level translation evaluators via MQM error span annotations. With modern LLMs supporting larger context windows, a natural question arises: can we feed entire document translations into an LLM for quality assessment? Ideally, evaluation should be invariant to text length, producing consistent error spans regardless of input granularity. However, our analysis shows that text length significantly impacts evaluation: longer texts lead to fewer error spans and reduced system ranking accuracy. To address this limitation, we evaluate several strategies, including granularity-aligned prompting, Focus Sentence Prompting (FSP), and a fine-tuning approach to better align LLMs with the evaluation task. The latter two methods largely mitigate this length bias, making LLMs more reliable for long-form translation evaluation.
PricingLogic: Evaluating LLMs Reasoning on Complex Tourism Pricing Tasks
Yunuo Liu | Dawei Zhu | Zena Al-Khalili | Dai Cheng | Yanjun Chen | Dietrich Klakow | Wei Zhang | Xiaoyu Shen
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Yunuo Liu | Dawei Zhu | Zena Al-Khalili | Dai Cheng | Yanjun Chen | Dietrich Klakow | Wei Zhang | Xiaoyu Shen
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
We present PricingLogic, the first benchmarkthat probes whether Large Language Mod-els (LLMs) can reliably automate tourism-booking prices when multiple, overlapping farerules apply. Travel agencies are eager to of-fload this error-prone task to AI systems; how-ever, deploying LLMs without verified reliabil-ity could result in significant financial lossesand erode customer trust. PricingLogic com-prises 300 natural-language questions based onbooking requests derived from 42 real-worldpricing policies, spanning two levels of diffi-culty: (i) basic customer-type pricing and (ii)bundled-tour calculations involving interactingdiscounts. Evaluations of a line of LLMs re-veal a steep performance drop on the harder tier,exposing systematic failures in rule interpreta-tion and arithmetic reasoning. These resultshighlight that, despite their general capabilities,today’s LLMs remain unreliable for revenue-critical applications without further safeguardsor domain adaptation. Our code and dataset areavaliable in https://github.com/EIT-NLP/PricingLogic.
AFRIDOC-MT: Document-level MT Corpus for African Languages
Jesujoba Oluwadara Alabi | Israel Abebe Azime | Miaoran Zhang | Cristina España-Bonet | Rachel Bawden | Dawei Zhu | David Ifeoluwa Adelani | Clement Oyeleke Odoje | Idris Akinade | Iffat Maab | Davis David | Shamsuddeen Hassan Muhammad | Neo Putini | David O. Ademuyiwa | Andrew Caines | Dietrich Klakow
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
Jesujoba Oluwadara Alabi | Israel Abebe Azime | Miaoran Zhang | Cristina España-Bonet | Rachel Bawden | Dawei Zhu | David Ifeoluwa Adelani | Clement Oyeleke Odoje | Idris Akinade | Iffat Maab | Davis David | Shamsuddeen Hassan Muhammad | Neo Putini | David O. Ademuyiwa | Andrew Caines | Dietrich Klakow
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
This paper introduces AFRIDOC-MT, a document-level multi-parallel translation dataset covering English and five African languages: Amharic, Hausa, Swahili, Yorùbá, and Zulu. The dataset comprises 334 health and 271 information technology news documents, all human-translated from English to these languages. We conduct document-level translation benchmark experiments by evaluating the ability of neural machine translation (NMT) models and large language models (LLMs) to translate between English and these languages, at both the sentence and pseudo-document levels, the outputs being realigned to form complete documents for evaluation. Our results indicate that NLLB-200 achieves the best average performance among the standard NMT models, while GPT-4o outperforms general-purpose LLMs. Fine-tuning selected models leads to substantial performance gains, but models trained on sentences struggle to generalize effectively to longer documents. Furthermore, our analysis reveals that some LLMs exhibit issues such as under-generation, over-generation, repetition of words and phrases, and off-target translations, specifically for translation into African languages.
Search
Fix author
Co-authors
- Xiaoyu Shen 3
- Yanjun Chen 2
- Dietrich Klakow 2
- David Ifeoluwa Adelani 1
- David O. Ademuyiwa 1
- Idris Akinade 1
- Zena Al-Khalili 1
- Jesujoba Alabi 1
- Israel Abebe Azime 1
- Matthias Aßenmacher 1
- Rachel Bawden 1
- Bill Byrne 1
- Andrew Caines 1
- Xinghao Chen 1
- Dai Cheng 1
- Davis David 1
- Michael Denkowski 1
- Tobias Domhan 1
- Yupei Du 1
- Cristina España-Bonet 1
- Michael Hahn 1
- Felix Hieber 1
- Alexander Koller 1
- Yunuo Liu 1
- Iffat Maab 1
- Shamsuddeen Hassan Muhammad 1
- Clement Oyeleke Odoje 1
- Neo Putini 1
- Leonardo F. R. Ribeiro 1
- Benjamin Roth 1
- Andreas Stephan 1
- Yirong Sun 1
- Shaomu Tan 1
- Ke Tran 1
- Sony Trenous 1
- Erjia Xiao 1
- Yuekun Yao 1
- Miaoran Zhang 1
- Wei Zhang 1