Nádia Félix Felipe Da Silva
Author directoryAlso published as: Nádia da Silva, Nádia F. F. da Silva, Nádia Da Silva, Nadia Felix Felipe da Silva, Nádia F. F. da Silva, Nádia Félix Felipe da Silva, Nádia Félix Felipe da Silva
2026
UFG-Semantic at SemEval-2026 Task 6: CLARITY - Unmasking Political Question Evasions
Aline Hamano | Beatriz Felicio | Henrique Galvão | Nádia Da Silva
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
Aline Hamano | Beatriz Felicio | Henrique Galvão | Nádia Da Silva
Proceedings of the 20th International Workshop on Semantic Evaluation (2026)
We propose an approach for Task 6: CLARITY - Unmasking Political Question Evasions. We make use of data augmentation, supervised fine-tuning, and model benchmarking to detect and classify response ambiguity in political discourse. Building on well-founded theory on equivocation and leveraging recent advancements in language modeling, our system was structured based on question/answer (QA) pairs extracted from presidential interviews, and it was evaluated in Clarity-level Classification and Evasion-level Classification.
LexIris-pt and LexBert-pt: Specialized Sentence Embeddings for Legal Similarity in Brazilian Portuguese
Willgnner Ferreira Santos | João Gabriel Grandotto Viana | Antônio Pires de Castro Júnior | Fernando Ribeiro Trindade | Nádia Félix Felipe da Silva
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Willgnner Ferreira Santos | João Gabriel Grandotto Viana | Antônio Pires de Castro Júnior | Fernando Ribeiro Trindade | Nádia Félix Felipe da Silva
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
This work presents and evaluates two specialized sentence embedding models for the Portuguese legal domain, LexIris-pt and LexBert-pt, obtained through supervised fine-tuning of BERT-based models using pairs of initial petitions. We propose a comparative evaluation protocol along three fronts: (i) zero-shot inference with pretrained embeddings, (ii) supervised fine-tuning on these pairs, and (iii) vector retrieval with incremental clustering over a corpus of 20,000 initial petitions. The results show that fine-tuning consistently increases correlations with reference scores and improves performance in vector retrieval; additionally, the vector retrieval stage indicates that the metric configured in the index (cosine similarity or inner product) can change the granularity of the partitioning under a fixed threshold, reinforcing the need for joint calibration among the encoder, metric and threshold. After auditing by specialists from the partner institution, LexIris-pt and LexBert-pt were operationally adopted to support the screening and organization of repetitive claims and predatory litigation.
BIPA: Brazilian Portuguese Phonetic Dataset with Dialectal Variations in IPA Standard
Thiago Monteles de Sousa | Lucas Rafael Gris | Nádia Félix Felipe da Silva
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Thiago Monteles de Sousa | Lucas Rafael Gris | Nádia Félix Felipe da Silva
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
This work presents BIPA, a phonetic transcription corpus for Brazilian Portuguese that covers regional dialectal variations. The corpus was constructed through automated extraction from Wiktionary, resulting in 53,353 unique words and 350,021 transcriptions in IPA format, distributed across six dialects: general Brazilian, Rio de Janeiro, São Paulo, South Region, Northeast Region, and Center-West Region. The average density of 6.56 transcriptions per word reflects multiple regionally conditioned phonetic variations. To validate the utility of the corpus, the ByT5-small model was fine-tuned for grapheme-to-phoneme conversion, achieving a Minimum Phoneme Error Rate of 2.66% on the validation set. BIPA addresses the scarcity of computational linguistic resources for Brazilian Portuguese, enabling applications in regional speech synthesis, automatic accent recognition, and computational sociolinguistic analysis.
The PROPOR Ecosystem: Structure, Roles, and Evolution of Portuguese-Language NLP
Rafael O. Nunes | Gustavo L. Tamiosso | Pedro L. C. de Andrade | Matheus S. de Aguiar | Rafael P. de Gouveia | Higor Moreira | Bruno Tavares | Laura P. de Gouveia | Felipe S. F. Paula | Andre Spritzer | Hidelberg O. Albuquerque | Nádia F. F. da Silva | Ellen P. R. S. Pereira | Dennis G. Balreira | Joel L. Carbonera
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Rafael O. Nunes | Gustavo L. Tamiosso | Pedro L. C. de Andrade | Matheus S. de Aguiar | Rafael P. de Gouveia | Higor Moreira | Bruno Tavares | Laura P. de Gouveia | Felipe S. F. Paula | Andre Spritzer | Hidelberg O. Albuquerque | Nádia F. F. da Silva | Ellen P. R. S. Pereira | Dennis G. Balreira | Joel L. Carbonera
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
The PROPOR conference has been the main venue for Portuguese language Natural Language Processing (NLP) research for over two decades. This paper presents a longitudinal bibliometric analysis of PROPOR from 2003 to 2024, examining thematic evolution, community structure, and scientific impact. We identify a shift from speech-oriented research toward text-based tasks, alongside the sustained importance of resources and linguistic theory. The community exhibits a stable structure, with complementary leadership models centered on institutional hubs and brokerage roles. Scientific impact is highly concentrated, following a long tail distribution, and distinguishes between cumulative productivity-driven impact and rapidly accelerating citation uptake in recent editions. These findings characterize PROPOR as a resilient regional linguistic ecosystem evolving in dialogue with broader NLP paradigms.
UlyssesLegalNER-Br: from Legislative to Legal, a comprehensive corpus of Brazilian legal documents for Named Entity Recognition
Hidelberg O. Albuquerque | Ellen Souza | Danilo C. G. Lucena | Héldon J. O. Albuquerque | Nádia F. F. da Silva | Márcio de S. Dias | Rafael O. Nunes | Adriano L. I. Oliveira | André C. P. L. F. de Carvalho
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
Hidelberg O. Albuquerque | Ellen Souza | Danilo C. G. Lucena | Héldon J. O. Albuquerque | Nádia F. F. da Silva | Márcio de S. Dias | Rafael O. Nunes | Adriano L. I. Oliveira | André C. P. L. F. de Carvalho
Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1
The legal domain presents several challenges for Natural Language Processing (NLP), particularly due to its linguistic complexity and lack of public datasets. Named Entity Recognition (NER), a subarea of NLP, has been successfully used to extract useful knowledge from legal texts. Its widespread use is limited by the lack of legal text corpora. This paper introduces UlyssesLegalNER-Br, a comprehensive corpus of Brazilian legal documents for NER, covering bills, case laws and laws, including the first NER corpus based exclusively on Brazilian laws. This research expand the UlyssesNER-Br corpus, previously focused only on the Brazilian legislative domain. The proposed corpus has 560 public documents annotated using a hybrid approach, organized in 9 categories and 23 fine-grained types, experimentally evaluated with the CRF, BiLSTM, and BERTimbau architectures. The corpus was experimentally evaluated regarding predictive performance, computational cost and label-level results. The best micro F1 96.18% was achieved by BERTimbau on the unified corpus, providing a strong baseline for Brazilian legal NER. At the label level, six categories and seven types presented a F1-score above 95%, while the lowest were distributed in the interval 71-82%.
2025
Labor Lex: A New Portuguese Corpus and Pipeline for Information Extraction in Brazilian Legal Texts
Pedro Vitor Quinta de Castro | Nádia Félix Felipe Da Silva
Proceedings of the Natural Legal Language Processing Workshop 2025
Pedro Vitor Quinta de Castro | Nádia Félix Felipe Da Silva
Proceedings of the Natural Legal Language Processing Workshop 2025
Relation Extraction (RE) is a challenging Natural Language Processing task that involves identifying named entities from text and classifying the relationships between them. When applied to a specific domain, the task acquires a new layer of complexity, handling the lexicon and context particular to the domain in question. In this work, this task is applied to the Legal domain, specifically targeting Brazilian Labor Law. Architectures based on Deep Learning, with word representations derived from Transformer Language Models (LM), have shown state-of-the-art performance for the RE task. Recent works on this task handle Named Entity Recognition (NER) and RE either as a single joint model or as a pipelined approach. In this work, we introduce Labor Lex, a newly constructed corpus based on public documents from Brazilian Labor Courts. We also present a pipeline of models trained on it. Different experiments are conducted for each task, comparing supervised training using LMs and In-Context Learning (ICL) with Large Language Models (LLM), and verifying and analyzing the results for each one. For the NER task, the best achieved result was 89.97% F1-Score, and for the RE task, the best result was 82.38% F1-Score. The best results for both tasks were obtained using the supervised training approach.
2024
Natural Language Processing Application in Legislative Activity: a Case Study of Similar Amendments in the Brazilian Senate
Diany Pressato | Pedro L. C. de Andrade | Flávio R. Junior | Felipe A. Siqueira | Ellen Polliana R. Souza | Nádia F. F. da Silva | Márcio de S. Dias | André C. P. L. F. de Carvalho
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Diany Pressato | Pedro L. C. de Andrade | Flávio R. Junior | Felipe A. Siqueira | Ellen Polliana R. Souza | Nádia F. F. da Silva | Márcio de S. Dias | André C. P. L. F. de Carvalho
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Enhancing Stance Detection in Low-Resource Brazilian Portuguese Using Corpus Expansion generated by GPT-3.5
Dyonnatan Maia | Nádia Félix Felipe da Silva
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Dyonnatan Maia | Nádia Félix Felipe da Silva
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
UlyssesNERQ: Expanding Queries from Brazilian Portuguese Legislative Documents through Named Entity Recognition
Hidelberg O. Albuquerque | Ellen Souza | Tainan Silva | Rafael P. Gouveia | Flavio Junior | Douglas Vitório | Nádia F. F. da Silva | André C.P.L.F. de Carvalho | Adriano L.I. Oliveira | Francisco Edmundo de Andrade
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
Hidelberg O. Albuquerque | Ellen Souza | Tainan Silva | Rafael P. Gouveia | Flavio Junior | Douglas Vitório | Nádia F. F. da Silva | André C.P.L.F. de Carvalho | Adriano L.I. Oliveira | Francisco Edmundo de Andrade
Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1
2023
Tecnicas de sumarização de textos juridicos para suporte à classificação de documentos de decisões judiciais
Hellen Harada | Fabiola Pereira | Alex Almeida | Daniela Freire | Marcio Dias | Nadia Felix Felipe da Silva | Pedro Andrade | Andre Carvalho
Proceedings of the 14th Brazilian Symposium in Information and Human Language Technology
Hellen Harada | Fabiola Pereira | Alex Almeida | Daniela Freire | Marcio Dias | Nadia Felix Felipe da Silva | Pedro Andrade | Andre Carvalho
Proceedings of the 14th Brazilian Symposium in Information and Human Language Technology
CAPIVARA: Cost-Efficient Approach for Improving Multilingual CLIP Performance on Low-Resource Languages
Gabriel Oliveira dos Santos | Diego Alysson Braga Moreira | Alef Iury Ferreira | Jhessica Silva | Luiz Pereira | Pedro Bueno | Thiago Sousa | Helena Maia | Nádia Da Silva | Esther Colombini | Helio Pedrini | Sandra Avila
Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)
Gabriel Oliveira dos Santos | Diego Alysson Braga Moreira | Alef Iury Ferreira | Jhessica Silva | Luiz Pereira | Pedro Bueno | Thiago Sousa | Helena Maia | Nádia Da Silva | Esther Colombini | Helio Pedrini | Sandra Avila
Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)
DeepLearningBrasil@LT-EDI-2023: Exploring Deep Learning Techniques for Detecting Depression in Social Media Text
Eduardo Garcia | Juliana Gomes | Adalberto Barbosa Junior | Cardeque Borges | Nádia da Silva
Proceedings of the Third Workshop on Language Technology for Equality, Diversity and Inclusion
Eduardo Garcia | Juliana Gomes | Adalberto Barbosa Junior | Cardeque Borges | Nádia da Silva
Proceedings of the Third Workshop on Language Technology for Equality, Diversity and Inclusion
In this paper, we delineate the strategy employed by our team, DeepLearningBrasil, which secured us the first place in the shared task DepSign-LT-EDI@RANLP-2023 with the advantage of 2.4%. The task was to classify social media texts into three distinct levels of depression - “not depressed,” “moderately depressed,” and “severely depressed.” Leveraging the power of the RoBERTa and DeBERTa models, we further pre-trained them on a collected Reddit dataset, specifically curated from mental health-related Reddit’s communities (Subreddits), leading to an enhanced understanding of nuanced mental health discourse. To address lengthy textual data, we introduced truncation techniques that retained the essence of the content by focusing on its beginnings and endings. Our model was robust against unbalanced data by incorporating sample weights into the loss. Cross-validation and ensemble techniques were then employed to combine our k-fold trained models, delivering an optimal solution. The accompanying code is made available for transparency and further development.
Search
Fix author
Co-authors
- Hidelberg O. Albuquerque 3
- André C. P. L. F. de Carvalho 3
- Pedro L. C. de Andrade 2
- Márcio de S. Dias 2
- Rafael O. Nunes 2
- Adriano L. I. Oliveira 2
- Ellen Souza 2
- Matheus S. de Aguiar 1
- Héldon J. O. Albuquerque 1
- Alex Almeida 1
- Pedro Andrade 1
- Sandra Avila 1
- Dennis G. Balreira 1
- Cardeque Borges 1
- Diego Alysson Braga Moreira 1
- Pedro Bueno 1
- Joel L. Carbonera 1
- André Carvalho 1
- Esther Colombini 1
- Márcio Dias 1
- Beatriz Felicio 1
- Alef Iury Ferreira 1
- Daniela Freire 1
- Henrique Galvão 1
- Eduardo Garcia 1
- Juliana Gomes 1
- Laura P. de Gouveia 1
- Rafael P. Gouveia 1
- Rafael P. de Gouveia 1
- Lucas Rafael Gris 1
- Aline Hamano 1
- Hellen Harada 1
- Adalberto Barbosa Junior 1
- Flavio Junior 1
- Flávio R. Junior 1
- Antônio Pires de Castro Júnior 1
- Danilo C. G. Lucena 1
- Dyonnatan Maia 1
- Helena Maia 1
- Higor Moreira 1
- Felipe S. F. Paula 1
- Hélio Pedrini 1
- Ellen P. R. S. Pereira 1
- Fabiola Pereira 1
- Luiz Pereira 1
- Diany Pressato 1
- Willgnner Ferreira Santos 1
- Jhessica Silva 1
- Tainan Silva 1
- Felipe A. Siqueira 1
- Thiago Sousa 1
- Thiago Monteles de Sousa 1
- Ellen Polliana R. Souza 1
- André Spritzer 1
- Gustavo Lopes Tamiosso 1
- Bruno Tavares 1
- Fernando Ribeiro Trindade 1
- João Gabriel Grandotto Viana 1
- Douglas Vitório 1
- Francisco Edmundo de Andrade 1
- Pedro Vitor Quinta de Castro 1
- Gabriel Oliveira dos Santos 1