Sergi Alvarez-Vidal
Author directoryAlso published as: Sergi Alvarez Vidal, Sergi Álvarez Vidal, Sergi Àlvarez Vidal, Sergi Álvarez, Sergi Alvarez, Sergi Álvarez-Vidal
2026
Teaching Machine Translation Technologies with MTUUOC
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026)
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026)
This paper presents the pedagogical integration of MTUOC, an open-source project developed at Universitat Oberta de Catalunya (UOC)—a distance-learning institution—to facilitate the training, fine-tuning, and integration of Neural Machine Translation (NMT) and Large Language Models (LLMs). The project consists of a modular suite of tools designed to streamline complex technical workflows for translation purposes. These components are currently utilised across research, industry knowledge transfer, and formal education. Specifically, the tools have been successfully implemented in a Bachelor’s degree in Translation and Interpreting and a Master’s degree in Translation Technologies. Furthermore, a pilot open course based on this framework received significant interest, reaching over 100 participants. This paper outlines the core components of the project, discusses the teaching experiences gathered in asynchronous environments, and describes the organisation of a forthcoming open course scheduled for October 2026. The results suggest that providing students with accessible, high-level interfaces for AI-based translation technologies enhances their technical autonomy and professional readiness.
Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026)
Ralph Krüger | Dorothy Kenny | Sheila Castilho | Sergi Álvarez-Vidal | Nora Aranberri | María Isabel Rivas Ginel | Janiça Hackenbuchner
Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026)
Ralph Krüger | Dorothy Kenny | Sheila Castilho | Sergi Álvarez-Vidal | Nora Aranberri | María Isabel Rivas Ginel | Janiça Hackenbuchner
Proceedings of the 1st International Workshop on Teaching AI-Based Translation and Technologies (TAITT 2026)
TaMTAS: Terminology-Aware Machine Translation for Accessible Science. Large Corpus compilation, terminology extraction and data augmentation
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Antoni Oliver | Sergi Alvarez-Vidal
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
This paper presents the TaMTAS project (Terminology-Aware Machine Translation for Accessible Science), a research project coordinated by the Universitat Oberta de Catalunya (UOC) to develop an open- source translation ecosystem for the Life Sciences. While we provide a general overview of the project’s organization into seven Work Packages (WPs) and its col- laborative consortium, this article focuses specifically on the work of WP2. Led by the UOC, this package is responsible for the parallel corpus compilation for five lan- guages (English, Spanish, Catalan, Esto- nian, and Irish), the enhancement of TBX- Tools for terminology extraction, and the development of synthetic data augmenta- tion strategies. These linguistic assets are essential to power the downstream Large Reasoning Models (LRMs) and Automatic Post-Editing (APE) modules, ensuring ter- minological consistency in highly special- ized scientific domains.
Literacy-Grounded and Industry-Oriented Translation Training with LT-LiDER
Janiça Hackenbuchner | María Isabel Rivas Ginel | Joss Moorkens | Sheila Castilho | Nora Aranberri | Sergi Álvarez Vidal | María do Campo Bayón | Ralph Krüger
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
Janiça Hackenbuchner | María Isabel Rivas Ginel | Joss Moorkens | Sheila Castilho | Nora Aranberri | Sergi Álvarez Vidal | María do Campo Bayón | Ralph Krüger
Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2)
he Erasmus+-funded international research consortium LT-LiDER develops a range of digital training resources which are grounded in the overarching frameworks of digital and AI literacy and oriented towards practical application contexts in the language and translation industry. These resources can be implemented on a component basis or as a complete curriculum in higher-education language and translation classrooms.
A Comparative Study in Corpus Linguistics Applied to Automatic Terminology Extraction
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Mercè Vàzquez | Sergi Alvarez-Vidal | Antoni Oliver
Proceedings of the 19th Workshop on Building and Using Comparable Corpora (BUCC)
Parallel and comparable corpora are the main linguistic resources to identify multilingual terminology using automatic term extraction tools. However, parallel corpora are available only for certain languages, domains and genres, and comparable corpora have some limitations when identifying corresponding terms. To implement a more efficient selection of multilingual terminology, we compared the performance of using specialised parallel and comparable corpora applied to languages with various forms of capital in linguistic resources. This paper presents a comparative study in corpus linguistics in which we automatically identify terms in Catalan, Spanish and English in legislation and administrative law using parallel corpora, comparable corpora and a combined methodology based on both typologies of corpora together with word embeddings. We observe that the combined methodology implemented obtains a higher number of term candidates than when working exclusively with parallel or comparable corpora. The evaluation of the results is performed using a terminological thesaurus as a gold standard. The new methodology presented in our study permits us to identify multilingual terminology in an efficient way, especially in Catalan-Spanish languages.
2025
Using Translation Techniques to Characterize MT Outputs
Sergi Alvarez-Vidal | Maria de Campo | Christian Olalla-Soler | Pilar Sánchez-Gijón
Proceedings of Machine Translation Summit XX: Volume 1
Sergi Alvarez-Vidal | Maria de Campo | Christian Olalla-Soler | Pilar Sánchez-Gijón
Proceedings of Machine Translation Summit XX: Volume 1
While current NMT and GPT models improve fluency and context awareness, they struggle with creative texts, where figurative language and stylistic choices are crucial. Current evaluation methods fail to capture these nuances, which requires a more descriptive approach. We propose a taxonomy based on translation techniques to assess machine-generated translations more comprehensively. The pilot study we conducted comparing human machine-produced translations reveals that human translations employ a wider range of techniques, enhancing naturalness and cultural adaptation. NMT and GPT models, even with prompting, tend to simplify content and introduce accuracy errors. Our findings highlight the need for refined frameworks that consider stylistic and contextual accuracy, ultimately bridging the gap between human and machine translation performance.
Fine-tuning and evaluation of NMT models for literary texts using RomCro v.2.0
Bojana Mikelenić | Antoni Oliver | Sergi Àlvarez Vidal
Proceedings of the Second Workshop on Creative-text Translation and Technology (CTT)
Bojana Mikelenić | Antoni Oliver | Sergi Àlvarez Vidal
Proceedings of the Second Workshop on Creative-text Translation and Technology (CTT)
This paper explores the fine-tuning and evaluation of neural machine translation (NMT) models for literary texts using RomCro v.2.0, an expanded multilingual and multidirectional parallel corpus. RomCro v.2.0 is based on RomCro v.1.0, but includes additional literary works, as well as texts in Catalan, making it a valuable resource for improving MT in underrepresented language pairs. Given the challenges of literary translation, where style, narrative voice, and cultural nuances must be preserved, fine-tuning on high-quality domain-specific data is essential for enhancing MT performance. We fine-tune existing NMT models with RomCro v.2.0 and evaluate their performance for six different language combinations using automatic metrics and for Spanish-Croatian and French-Catalan using manual evaluation. Results indicate that fine-tuned models outperform general-purpose systems, achieving greater fluency and stylistic coherence. These findings support the effectiveness of corpus-driven fine-tuning for literary translation and highlight the importance of curated high-quality corpus.
2024
Training an NMT system for legal texts of a low-resource language variety South Tyrolean German - Italian
Antoni Oliver | Sergi Álvarez | Egon W. Stemle | Elena Chiocchetti
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
Antoni Oliver | Sergi Álvarez | Egon W. Stemle | Elena Chiocchetti
Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1)
This paper illustrates the process of training and evaluating NMT systems for a language pair that includes a low-resource language variety.A parallel corpus of legal texts for Italian and South Tyrolean German has been compiled, with South Tyrolean German being the low-resourced language variety. As the size of the compiled corpus is insufficient for the training, we have combined the corpus with several parallel corpora using data weighting at sentence level. We then performed an evaluation of each combination and of two popular commercial systems.
LitPC: A set of tools for building parallel corpora from literary works
Antoni Oliver | Sergi Álvarez
Proceedings of the 1st Workshop on Creative-text Translation and Technology
Antoni Oliver | Sergi Álvarez
Proceedings of the 1st Workshop on Creative-text Translation and Technology
In this paper, we describe the LitPC toolkit, a variety of tools and methods designed for the quick and effective creation of parallel corpora derived from literary works. This toolkit can be a useful resource due to the scarcity of curated parallel texts for this domain. We also feature a case study describing the creation of a Russian-English parallel corpus based on the literary works by Leo Tolstoy. Furthermore, an augmented version of this corpus is used to both train and assess neural machine translation systems specifically adapted to the author’s style.
2023
TAN-IBE: Neural Machine Translation for the romance languages of the Iberian Peninsula
Antoni Oliver | Mercè Vàzquez | Marta Coll-Florit | Sergi Álvarez | Víctor Suárez | Claudi Aventín-Boya | Cristina Valdés | Mar Font | Alejandro Pardos
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Antoni Oliver | Mercè Vàzquez | Marta Coll-Florit | Sergi Álvarez | Víctor Suárez | Claudi Aventín-Boya | Cristina Valdés | Mar Font | Alejandro Pardos
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
The main goal of this project is to explore the techniques for training NMT systems applied to Spanish, Portuguese, Catalan, Galician, Asturian, Aragonese and Aranese. These languages belong to the same Romance family, but they are very different in terms of the linguistic resources available. Asturian, Aragonese and Aranese can be considered low resource languages. These characteristics make this setting an excellent place to explore training techniques for low-resource languages: transfer learning and multilingual systems, among others. The first months of the project have been dedicated to the compilation of monolingual and parallel corpora for Asturian, Aragonese and Aranese.
Filtering and rescoring the CCMatrix corpus for Neural Machine Translation training
Antoni Oliver González | Sergi Álvarez
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Antoni Oliver González | Sergi Álvarez
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
There are several parallel corpora available for many language pairs, such as CCMatrix, built from mass downloads of web content and automatic detection of segments in one language and the translation equivalent in another. These techniques can produce large parallel corpora, but of questionable quality. In many cases, the segments are not in the required languages, or if they are, they are not translation equivalents. In this article, we present an algorithm for filtering out the segments in languages other than the required ones and re-scoring the segments using SBERT. A use case on the Spanish-Asturian and Spanish-Catalan CCMatrix corpus is presented.
PE effort and neural-based automatic MT metrics: do they correlate?
Sergi Alvarez | Antoni Oliver
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Sergi Alvarez | Antoni Oliver
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Neural machine translation (NMT) has shown overwhelmingly good results in recent times. This improvement in quality has boosted the presence of NMT in nearly all fields of translation. Most current translation industry workflows include postediting (PE) of MT as part of their process. For many domains and language combinations, translators post-edit raw machine translation (MT) to produce the final document. However, this process can only work properly if the quality of the raw MT output can be assured. MT is usually evaluated using automatic scores, as they are much faster and cheaper. However, traditional automatic scores have not been good quality indicators and do not correlate with PE effort. We analyze the correlation of each of the three dimensions of PE effort (temporal, technical and cognitive) with COMET, a neural framework which has obtained outstanding results in recent MT evaluation campaigns.
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Mary Nurminen | Judith Brenner | Maarit Koponen | Sirkku Latomaa | Mikhail Mikhailov | Frederike Schierl | Tharindu Ranasinghe | Eva Vanmassenhove | Sergi Alvarez Vidal | Nora Aranberri | Mara Nunziatini | Carla Parra Escartín | Mikel Forcada | Maja Popovic | Carolina Scarton | Helena Moniz
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Mary Nurminen | Judith Brenner | Maarit Koponen | Sirkku Latomaa | Mikhail Mikhailov | Frederike Schierl | Tharindu Ranasinghe | Eva Vanmassenhove | Sergi Alvarez Vidal | Nora Aranberri | Mara Nunziatini | Carla Parra Escartín | Mikel Forcada | Maja Popovic | Carolina Scarton | Helena Moniz
Proceedings of the 24th Annual Conference of the European Association for Machine Translation
Training and integration of neural machine translation with MTUOC
Antoni Oliver | Sergi Alvarez
Proceedings of the 1st Workshop on Open Community-Driven Machine Translation
Antoni Oliver | Sergi Alvarez
Proceedings of the 1st Workshop on Open Community-Driven Machine Translation
In this paper the goals and main objectives of the project MTUOC are presented. This project aims to ease the process of training and integrating neural machine translation (NMT) systems into professional translation environments. The MTUOC project distributes a series of auxiliary tools that allow to perform parallel corpus compilation and preprocessing, as well as the training of NMT systems. The project also distributes a server that implements most of the communication protocols used in computer assisted translation tools.
2020
Quantitative Analysis of Post-Editing Effort Indicators for NMT
Sergi Alvarez | Antoni Oliver | Toni Badia
Proceedings of the 22nd Annual Conference of the European Association for Machine Translation
Sergi Alvarez | Antoni Oliver | Toni Badia
Proceedings of the 22nd Annual Conference of the European Association for Machine Translation
The recent improvements in machine translation (MT) have boosted the use of post-editing (PE) in the translation industry. A new machine translation paradigm, neural machine translation (NMT), is displacing its corpus-based predecessor, statistical machine translation (SMT), in the translation workflows currently implemented because it usually increases the fluency and accuracy of the MT output. However, usual automatic measurements do not always indicate the quality of the MT output and there is still no clear correlation between PE effort and productivity. We present a quantitative analysis of different PE effort indicators for two NMT systems (transformer and seq2seq) for English-Spanish in-domain medical documents. We compare both systems and study the correlation between PE time and other scores. Results show less PE effort for the transformer NMT model and a high correlation between PE time and keystrokes.
PosEdiOn: Post-Editing Assessment in PythOn
Antoni Oliver | Sergi Alvarez | Toni Badia
Proceedings of the 22nd Annual Conference of the European Association for Machine Translation
Antoni Oliver | Sergi Alvarez | Toni Badia
Proceedings of the 22nd Annual Conference of the European Association for Machine Translation
There is currently an extended use of post-editing of machine translation (PEMT) in the translation industry. This is due to the increase in the demand of translation and to the significant improvements in quality achieved by neural machine translation (NMT). PEMT has been included as part of the translation workflow because it increases translators’ productivity and it also reduces costs. Although an effective post-editing requires enough quality of the MT output, usual automatic metrics do not always correlate with post-editing effort. We describe a standalone tool designed both for industry and research that has two main purposes: collect sentence-level information from the post-editing process (e.g. post-editing time and keystrokes) and visually present multiple evaluation scores so they can be easily interpreted by a user.
2019
Search
Fix author
Co-authors
- Antoni Oliver 12
- Nora Aranberri 3
- Toni Badia 3
- Sheila Castilho 2
- María Isabel Rivas Ginel 2
- Janiça Hackenbuchner 2
- Ralph Krüger 2
- Mercè Vàzquez 2
- Claudi Aventín-Boya 1
- María Do Campo Bayón 1
- Judith Brenner 1
- Elena Chiocchetti 1
- Marta Coll-Florit 1
- Mar Font 1
- Mikel L. Forcada 1
- Antoni Oliver González 1
- Dorothy Kenny 1
- Maarit Koponen 1
- Sirkku Latomaa 1
- Bojana Mikelenić 1
- Mikhail Mikhailov 1
- Helena Moniz 1
- Joss Moorkens 1
- Mara Nunziatini 1
- Mary Nurminen 1
- Christian Olalla-Soler 1
- Alejandro Pardos 1
- Carla Parra Escartín 1
- Maja Popović 1
- Tharindu Ranasinghe 1
- Carolina Scarton 1
- Frederike Schierl 1
- Egon Stemle 1
- Víctor Suárez 1
- Pilar Sánchez-Gijón 1
- Cristina Valdés 1
- Eva Vanmassenhove 1
- Maria de Campo 1