AnandaSky: A Vision–Language Model for Line-Level Transcription of Historical Sinographic Documents

Colin Brisson, Ayoub Kahfy, Frédéric Constant, Marc Bui


Abstract
We present AnandaSky, a vision–language model for line-level transcription of historical sinographic documents. The model combines a compact high-resolution visual encoder with global attention, 10px patches, uncompressed visual prefix and a Qwen3-0.6B autoregressive decoder. It is trained at scale on 4M annotated lines from documents produced in China and Korea between the 8th and 20th centuries. Across in-domain and held-out public benchmarks, AnandaSky achieves sub-1% CER on five of eight datasets, sets a new state of the art on MTHv2 with 0.92% CER, and shows strong transfer to unseen collections. For EvaHan 2026, full fine-tuning on the organizers’ data to match task-specific annotation conventions reduces CER relative to the official baseline by 5.2% on prints and 12.1% on manuscripts, despite using one-tenth as many parameters.
Anthology ID:
2026.lt4hala-1.32
Volume:
Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Rachele Sprugnoli, Marco Passarotti
Venues:
LT4HALA | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
311–321
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-lt4hala-32
DOI:
10.63317/3pk7cv8hxzod
Bibkey:
Cite (ACL):
Colin Brisson, Ayoub Kahfy, Frédéric Constant, and Marc Bui. 2026. AnandaSky: A Vision–Language Model for Line-Level Transcription of Historical Sinographic Documents. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA 2026) @ LREC 2026, pages 311–321, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
AnandaSky: A Vision–Language Model for Line-Level Transcription of Historical Sinographic Documents (Brisson et al., LT4HALA 2026)
Copy Citation: