Brenda Murphy
Author directory2026
BRIDGE-MT: A Benchmark for Role Interactions and Dependencies in Machine Translation Gender Evaluation
Neha Gajakos | Christopher Staff | Brenda Murphy | John D. Kelleher | Rejwanul Haque
Proceedings of the 17th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)
Neha Gajakos | Christopher Staff | Brenda Murphy | John D. Kelleher | Rejwanul Haque
Proceedings of the 17th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track)
This paper investigates gender behavior in Hindi–English machine translation (MT) within multi-entity settings, especially when two occupational roles appear within the same sentence. Existing benchmarks often focus on single-referenced entities, leaving cross-role dependencies largely unexplored. We define a taxonomy of thirteen role-gender configurations covering masculine (m), feminine (f), and neutral (n) assignments and introduce BRIDGE-MT, a manually created dataset of 351 Hindi–English sentence pairs (594 role-level instances) in order to evaluate dual-role interactions. We evaluated commercial MT systems and multilingual LLMs, and propose neutral-comparison asymmetry metrics and a conditional interaction metric to analyze cross-role dependencies. Our results show that explicitly gendered roles achieve higher F1 scores than neutral-labelled roles. We also observe a consistent position effect, where Role B (the second role) tends to have lower accuracy and greater gender asymmetry than Role A (the first role) across all evaluated systems. Conditional interaction analysis further indicates that the gender assigned to one role can influence the translation of the other. These findings highlight the importance of evaluating gender behavior in multi-entity settings to better understand interaction-driven asymmetries in MT.