Jonathan Ayebakuro Orama
Author directory2026
Alignment Quality Degradation Across the Parallel–Comparable Spectrum: A Comparative Analysis
Audrey Mash | Jonathan Ayebakuro Orama | Marc Juvillà Garcia | Maite Melero
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Audrey Mash | Jonathan Ayebakuro Orama | Marc Juvillà Garcia | Maite Melero
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 1)
Sentence-level alignment systems have been developed and evaluated primarily on parallel data, leaving their behaviour across the broader parallel-comparable spectrum of real web content poorly understood. We present a stratified empirical study of alignment quality for Catalan-English using 300 document pairs across three parallelism bands defined by mean-max LaBSE cosine similarity. We compare four systems: a hierarchical alignment pipeline (DocAlign), an ablation with paragraph pre-filtering disabled (DocAlign-NoFilter), the flat aligner Vecalign, and a flat LaBSE greedy baseline. Evaluation uses human-annotated sentence pairs and coverage-weighted quality. Quality degrades at different rates by system type: hierarchical systems maintain usable-pair rates ranging from 25% to 51% on comparable data while flat systems collapse to 2-7%. Paragraph pre-filtering reduces output volume on comparable data while raising pair quality relative to the unfiltered ablation. Vecalign is statistically indistinguishable from the greedy baseline at all parallelism levels, suggesting that LaBSE embedding discrimination is the binding constraint on flat alignment quality. Failure mode analysis of 550 low-rated pairs identifies topical mismatch as the dominant failure mode, with structural noise concentrated in flat systems.