Hélio Pedrini
Also published as: Helio Pedrini
2026
The In-Car Sign Language Corpus (ICSL): A Multi-Modal Resource for Constrained-Space Sign Language Recognition
Raviteja Boddu | Guilherme Vieira Leite | Joed Lopes da Silva | Ângelo Benetti | Isabela Barbieri | Natália de Melo Afonso | Thyago Santos | Hélio Pedrini | Felipe Venâncio Barbosa | José Mario De Martino | Munir Georges | Alessandro Zimmer
Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion
Raviteja Boddu | Guilherme Vieira Leite | Joed Lopes da Silva | Ângelo Benetti | Isabela Barbieri | Natália de Melo Afonso | Thyago Santos | Hélio Pedrini | Felipe Venâncio Barbosa | José Mario De Martino | Munir Georges | Alessandro Zimmer
Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion
This paper addresses the challenges of using sign language within shared mobility services, such as taxis, carpools, or ride-sharing platforms. The use of sign language recognition (SLR) in real-world, confined environments, specifically vehicle interiors remains largely unexplored. To motivate research in this area, we present the In-Car Sign Language (ICSL) dataset for Brazilian Sign Language (Libras), with the long-term goal of improving public transport accessibility for the Deaf and Hard-of-Hearing community. The dataset consists of: (1) high-precision laboratory motion capture (MoCap) data to establish an idealized linguistic baseline and (2) real-world multi-modal in-car recordings captured using a 2D camera and 3D Time-of-Flight sensors. The dataset provides a basis for comparative analyses between synthesized signing avatar animations and recorded real signing interpreter videos, which enable future research into robust “in-the-wild” SLR models and domain adaptation. We describe in detail the use cases, the setup, the data collection protocol, and the metadata structure of the corpus. In total, we recorded a multimodal dataset exceeding 1.5 million frames, comprising the synchronized multimodal streams described above featuring Libras users across various in-car scenarios. The corpus is provided with gloss annotation of lexical signs and non-lexical sign language elements specially designed to support the training and evaluation of deep neural networks for constrained space recognition. In-vehicle signing offers a technically significant example of a constrained, occluded, and non-frontal environment. While recognizing the diverse communication strategies already employed by the Deaf community, identifying automotive-specific limitations provides a useful stepping stone for research into enhancing in-car accessibility and passenger quality of life.
2023
CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages
Gabriel Oliveira dos Santos | Diego Alysson Braga Moreira | Alef Iury Ferreira | Jhessica Silva | Luiz Pereira | Pedro Bueno | Thiago Sousa | Helena Maia | Nádia Da Silva | Esther Colombini | Helio Pedrini | Sandra Avila
Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)
Gabriel Oliveira dos Santos | Diego Alysson Braga Moreira | Alef Iury Ferreira | Jhessica Silva | Luiz Pereira | Pedro Bueno | Thiago Sousa | Helena Maia | Nádia Da Silva | Esther Colombini | Helio Pedrini | Sandra Avila
Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)
2020
Lite Training Strategies for Portuguese-English and English-Portuguese Translation
Alexandre Lopes | Rodrigo Nogueira | Roberto Lotufo | Helio Pedrini
Proceedings of the Fifth Conference on Machine Translation
Alexandre Lopes | Rodrigo Nogueira | Roberto Lotufo | Helio Pedrini
Proceedings of the Fifth Conference on Machine Translation
Despite the widespread adoption of deep learning for machine translation, it is still expensive to develop high-quality translation models. In this work, we investigate the use of pre-trained models, such as T5 for Portuguese-English and English-Portuguese translation tasks using low-cost hardware. We explore the use of Portuguese and English pre-trained language models and propose an adaptation of the English tokenizer to represent Portuguese characters, such as diaeresis, acute and grave accents. We compare our models to the Google Translate API and MarianMT on a subset of the ParaCrawl dataset, as well as to the winning submission to the WMT19 Biomedical Translation Shared Task. We also describe our submission to the WMT20 Biomedical Translation Shared Task. Our results show that our models have a competitive performance to state-of-the-art models while being trained on modest hardware (a single 8GB gaming GPU for nine days). Our data, models and code are available in our GitHub repository.
Search
Fix author
Co-authors
- Sandra Avila 1
- Isabela Barbieri 1
- Ângelo Benetti 1
- Raviteja Boddu 1
- Diego Alysson Braga Moreira 1
- Pedro Bueno 1
- Esther Colombini 1
- Nádia Da Silva 1
- Alef Iury Ferreira 1
- Munir Georges 1
- Alexandre Lopes 1
- Joed Lopes da Silva 1
- Roberto Lotufo 1
- Helena Maia 1
- José Mario De Martino 1
- Rodrigo Nogueira 1
- Luiz Pereira 1
- Thyago Santos 1
- Jhessica Silva 1
- Thiago Sousa 1
- Felipe Venâncio Barbosa 1
- Guilherme Vieira Leite 1
- Alessandro Zimmer 1
- Natália de Melo Afonso 1
- Gabriel Oliveira dos Santos 1