ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation

Michał Ciesiółka, Dawid Wiśniewski, Adrian Charkiewicz, Kamil Guttmann


Abstract
We present ForMaT (Format-Preserving Multilingual Translation), a parallel corpus of 3,956 PDFs across 15 language pairs that preserves original layout metadata proposed for multimodal machine translation. To ensure structural diversity in the dataset, we employ K-Medoids sampling over 45 geometric features, capturing complex elements like nested tables and formulas to focus only on visually diverse PDF documents. Our evaluation reveals that current MT systems struggle with spatial grounding and geometric synchronization, often losing the link between text and its visual context. ForMaT provides a benchmark for developing layout-aware translation models that integrate visual and textual context for high-fidelity document reconstruction.
Anthology ID:
2026.eamt-1.12
Volume:
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Month:
June
Year:
2026
Address:
Tilburg, The Netherlands
Editors:
Dimitar Shterionov, Eva Vanmassenhove, Mirella De Sisto, Fred Blain, Javad Pourmostafa Roshan Sharami, Lisa Lepp, Chiara Manna, Argentina Anna Rescigno, Alina Karakanta, Ayla Rigouts Terryn, Manuel Lardelli, Natalia Resende, Elena Murgolo, Janiça Hackenbuchner, Anna Zaretskaya, Miquel Esplà-Gomis, Thierry Etchegoyhen, Dagmar Gromann, Rachel Bawden, Barry Haddow, Sara Szoc, Mikel Forcada, Helena Moniz
Venue:
EAMT
SIG:
Publisher:
European Association for Machine Translation
Note:
Pages:
143–157
Language:
URL:
https://aclanthology.org/2026.eamt-1.12/
DOI:
Bibkey:
Cite (ACL):
Michał Ciesiółka, Dawid Wiśniewski, Adrian Charkiewicz, and Kamil Guttmann. 2026. ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation. In Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1), pages 143–157, Tilburg, The Netherlands. European Association for Machine Translation.
Cite (Informal):
ForMaT: Dataset for Visually-Grounded Multilingual PDF Translation (Ciesiółka et al., EAMT 2026)
Copy Citation:
PDF:
https://aclanthology.org/2026.eamt-1.12.pdf