Gokhan Dogru

Author directory

2026

Drawing on 23 student projects from a fourth-year Machine Translation and Post-editing course, this paper examines how asking students to compare LLM and NMT outputs, interpret metric results, and justify a post-editing choice reveals their evaluative judgement. Students translated short specialised English Wikipedia texts into Catalan or Spanish, generated four system outputs, evaluated them using automatic metrics and human adequacy/fluency assessment, selected one output for post-editing, and justified their decision in written reports. The analysis combines descriptive counts from 23 projects with qualitative coding of the 22 cases sup-ported by written reports. Results show that students did not treat automatic metrics as final authority: final post-editing selections often diverged from metric rankings and were justified through adequacy, fluency, terminology, and expected post-editing effort. The study therefore does not compare systems under benchmark conditions; it analyses how students justified system choice within an au-thentic classroom assignment.
Search
Co-authors
    Venues
    Fix author