Multimodal Ancient Document Parsing: Technical Report for EvaHan2026 Competition

Liqi He, Qiwei Li, Ziye Yang, Zuchao Li


Abstract
We present the multimodal Optical Character Recognition (OCR) and layout analysis methods developed for the EvaHan 2026 competition. Our approach is built upon the Qwen2.5-VL-7B-Instruct architecture and integrates two core strategies: (1) a reinforcement learning alignment pipeline utilizing Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO) to explicitly mitigate hallucination and coordinate instability; and (2) a four-stage curriculum learning framework that synthesizes domain-specific historical artifacts to enhance open-modality generalization. Using this approach, we achieve competitive results, notably reaching a Character Error Rate (CER) of 0.0303 on printed texts (Task A) and 0.0552 on handwritten manuscripts (Task C), as well as an Average Intersection over Union (IoU) of 0.7638 on layout element analysis (Task B).
Anthology ID:
2026.lt4hala-1.33
Volume:
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Rachele Sprugnoli, Marco Passarotti
Venues:
LT4HALA | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
322–329
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-lt4hala-33
DOI:
10.63317/2cfum2ozgjrs
Bibkey:
Cite (ACL):
Liqi He, Qiwei Li, Ziye Yang, and Zuchao Li. 2026. Multimodal Ancient Document Parsing: Technical Report for EvaHan2026 Competition. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026, pages 322–329, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
Multimodal Ancient Document Parsing: Technical Report for EvaHan2026 Competition (He et al., LT4HALA 2026)
Copy Citation: