Ely E. Matos
2026
Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violence
Lívia Dutra | Arthur Lorenzi | Lais Berno | Franciany Campos | Karoline Biscardi | Kenneth Brown | Marcelo Viridiano | Frederico Belcavello | Ely E. Matos | Olivia Guaranha | Erik Santos | Sofia Reinach | Tiago Timponi Torrent
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Lívia Dutra | Arthur Lorenzi | Lais Berno | Franciany Campos | Karoline Biscardi | Kenneth Brown | Marcelo Viridiano | Frederico Belcavello | Ely E. Matos | Olivia Guaranha | Erik Santos | Sofia Reinach | Tiago Timponi Torrent
Proceedings of the Fifteenth Language Resources and Evaluation Conference
We introduce a methodology for the identification of notifiable events in the domain of healthcare. The methodology harnesses semantic frames to define fine-grained patterns and search them in unstructured data, namely, open-text fields in e-medical records. We apply the methodology to the problem of underreporting of gender-based violence (GBV) in e-medical records produced during patients’ visits to primary care units. A total of eight patterns are defined and searched on a corpus of 21 million sentences in Brazilian Portuguese extracted from e-SUS APS. The results are manually evaluated by linguists and the precision of each pattern measured. Our findings reveal that the methodology effectively identifies reports of violence with a precision of 0.726, confirming its robustness. Designed as a transparent, efficient, low-carbon, and language-agnostic pipeline, the approach can be easily adapted to other health surveillance contexts, contributing to the broader, ethical, and explainable use of NLP in public health systems.
Evaluating the Impact of LLM-Assisted Annotation in a Perspectivized Setting: The Case of FrameNet Annotation
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Frederico Belcavello | Ely E. Matos | Arthur Lorenzi | Lisandra Bonoto | Livia Pádua Ruiz | Luiz Fernando Pereira | Victor Herbst | Yulla Liquer Navarro | Helen de Andrade Abreu | Lívia Vicente Dutra | Tiago Timponi Torrent
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
The use of LLM-based applications as a means to accelerate and/or substitute human labor in the creation of language resources and datasets is a reality. Nonetheless, despite the potential of such tools for linguistic research, an evaluation of their performance and impact on the creation of annotated datasets, especially under a perspectivized approach to NLP, is still missing. This paper contributes to the reduction of this gap by reporting on an extensive evaluation of the (semi-)automatization of FrameNet-like semantic annotation by the use of an LLM-based semantic role labeler. The methodology employed compares annotation time, coverage, and diversity in three experimental settings: manual, automatic, and semi-automatic annotation. Results show that the hybrid, semi-automatic annotation setting leads to increased frame diversity and similar annotation coverage, when compared to the human-only setting, while the automatic setting performs considerably worse in all metrics, except for annotation time, which remains similar.
A Frame and Canvas-Based Perspective-Encoding Methodology for Multimodal Semantic Annotation of Classroom Settings
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
Claudia Ferraz | Ely E. Matos | Frederico Belcavello | Julia Gasparetto | Juliana de Oliveira | Janina Wildfeuer | Tiago Timponi Torrent
Proceedings of the 22nd Joint ACL - ISO Workshop on Interoperable Semantic Annotation and Representation (ISA-22) @ LREC 2026
We propose a methodology for the multimodal semantic annotation of classroom interactions that takes interactional canvases and semantic frames as its core analytical categories. The approach enables the systematic recording of semantic correlations among interactants, communicative modes, and material supports involved in situated meaning-making processes. The methodology encodes participant perspective by relying on the temporal alignment of multiple video recordings of the same instructional event captured from different viewpoints, allowing for the representation of how meaning construction unfolds across perceptual and interactional positions. To operationalize the proposal, we introduce an annotation tool that implements the scheme and supports the integration of multimodal data streams within a unified semantic representation framework. We conclude by discussing the limitations of the current proposal and the possibilities for extending it to other interactional settings.
2024
Framed Multi30K: A Frame-Based Multimodal-Multilingual Dataset
Marcelo Viridiano | Arthur Lorenzi | Tiago Timponi Torrent | Ely E. Matos | Adriana S. Pagano | Natália Sathler Sigiliano | Maucha Gamonal | Helen de Andrade Abreu | Lívia Vicente Dutra | Mairon Samagaio | Mariane Carvalho | Franciany Campos | Gabrielly Azalim | Bruna Mazzei | Mateus Fonseca de Oliveira | Ana Carolina Luz | Livia Padua Ruiz | Júlia Bellei | Amanda Pestana | Josiane Costa | Iasmin Rabelo | Anna Beatriz Silva | Raquel Roza | Mariana Souza Mota | Igor Oliveira | Márcio Henrique Pelegrino de Freitas
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Marcelo Viridiano | Arthur Lorenzi | Tiago Timponi Torrent | Ely E. Matos | Adriana S. Pagano | Natália Sathler Sigiliano | Maucha Gamonal | Helen de Andrade Abreu | Lívia Vicente Dutra | Mairon Samagaio | Mariane Carvalho | Franciany Campos | Gabrielly Azalim | Bruna Mazzei | Mateus Fonseca de Oliveira | Ana Carolina Luz | Livia Padua Ruiz | Júlia Bellei | Amanda Pestana | Josiane Costa | Iasmin Rabelo | Anna Beatriz Silva | Raquel Roza | Mariana Souza Mota | Igor Oliveira | Márcio Henrique Pelegrino de Freitas
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
This paper presents Framed Multi30K (FM30K), a novel frame-based Brazilian Portuguese multimodal-multilingual dataset which i) extends the Multi30K dataset (Elliot et al., 2016) with 158,915 original Brazilian Portuguese descriptions, and 30,104 Brazilian Portuguese translations from original English descriptions; and ii) adds 2,677,613 frame evocation labels to the 158,915 English descriptions and to the ones created for Brazilian Portuguese; (iii) extends the Flickr30k Entities dataset (Plummer et al., 2015) with 190,608 frames and Frame Elements correlations with the existing phrase-to-region correlations.
Frame2: A FrameNet-based Multimodal Dataset for Tackling Text-image Interactions in Video
Frederico Belcavello | Tiago Timponi Torrent | Ely E. Matos | Adriana S. Pagano | Maucha Gamonal | Natalia Sigiliano | Lívia Vicente Dutra | Helen de Andrade Abreu | Mairon Samagaio | Mariane Carvalho | Franciany Campos | Gabrielly Azalim | Bruna Mazzei | Mateus Fonseca de Oliveira | Ana Carolina Loçasso Luz | Lívia Pádua Ruiz | Júlia Bellei | Amanda Pestana | Josiane Costa | Iasmin Rabelo | Anna Beatriz Silva | Raquel Roza | Mariana Souza | Igor Oliveira
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Frederico Belcavello | Tiago Timponi Torrent | Ely E. Matos | Adriana S. Pagano | Maucha Gamonal | Natalia Sigiliano | Lívia Vicente Dutra | Helen de Andrade Abreu | Mairon Samagaio | Mariane Carvalho | Franciany Campos | Gabrielly Azalim | Bruna Mazzei | Mateus Fonseca de Oliveira | Ana Carolina Loçasso Luz | Lívia Pádua Ruiz | Júlia Bellei | Amanda Pestana | Josiane Costa | Iasmin Rabelo | Anna Beatriz Silva | Raquel Roza | Mariana Souza | Igor Oliveira
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
This paper presents the Frame2 dataset, a multimodal dataset built from a corpus of a Brazilian travel TV show annotated for FrameNet categories for both the text and image communicative modes. Frame2 comprises 230 minutes of video, which are correlated with 2,915 sentences either transcribing the audio spoken during the episodes or the subtitling segments of the show where the host conducts interviews in English. For this first release of the dataset, a total of 11,796 annotation sets for the sentences and 6,841 for the video are included. Each of the former includes a target lexical unit evoking a frame or one or more frame elements. For each video annotation, a bounding box in the image is correlated with a frame, a frame element and lexical unit evoking a frame in FrameNet.
MoCCA: A Model of Comparative Concepts for Aligning Constructicons
Arthur Lorenzi | Peter Ljunglöf | Ben Lyngfelt | Tiago Timponi Torrent | William Croft | Alexander Ziem | Nina Böbel | Linnéa Bäckström | Peter Uhrig | Ely E. Matos
Proceedings of the 20th Joint ACL - ISO Workshop on Interoperable Semantic Annotation @ LREC-COLING 2024
Arthur Lorenzi | Peter Ljunglöf | Ben Lyngfelt | Tiago Timponi Torrent | William Croft | Alexander Ziem | Nina Böbel | Linnéa Bäckström | Peter Uhrig | Ely E. Matos
Proceedings of the 20th Joint ACL - ISO Workshop on Interoperable Semantic Annotation @ LREC-COLING 2024
This paper presents MoCCA, a Model of Comparative Concepts for Aligning Constructicons under development by a consortium of research groups building Constructicons of different languages including Brazilian Portuguese, English, German and Swedish. The Constructicons will be aligned by using comparative concepts (CCs) providing language-neutral definitions of linguistic properties. The CCs are drawn from typological research on grammatical categories and constructions, and from FrameNet frames, organized in a conceptual network. Language-specific constructions are linked to the CCs in accordance with general principles. MoCCA is organized into files of two types: a largely static CC Database file and multiple Linking files containing relations between constructions in a Constructicon and the CCs. Tools are planned to facilitate visualization of the CC network and linking of constructions to the CCs. All files and guidelines will be versioned, and a mechanism is set up to report cases where a language-specific construction cannot be easily linked to existing CCs.
Search
Fix author
Co-authors
- Tiago Timponi Torrent 6
- Frederico Belcavello 4
- Arthur Lorenzi 4
- Helen de Andrade Abreu 3
- Franciany Campos 3
- Lívia Vicente Dutra 3
- Lívia Pádua Ruiz 3
- Gabrielly Azalim 2
- Júlia Bellei 2
- Mariane Carvalho 2
- Josiane Costa 2
- Maucha Gamonal 2
- Bruna Mazzei 2
- Igor Oliveira 2
- Adriana Silvina Pagano 2
- Amanda Pestana 2
- Iasmin Rabelo 2
- Raquel Roza 2
- Mairon Samagaio 2
- Anna Beatriz Silva 2
- Marcelo Viridiano 2
- Mateus Fonseca de Oliveira 2
- Lais Berno 1
- Karoline Biscardi 1
- Lisandra Bonoto 1
- Kenneth Brown 1
- Linnéa Bäckström 1
- Nina Böbel 1
- William Croft 1
- Lívia Dutra 1
- Claudia Ferraz 1
- Julia Gasparetto 1
- Olívia Guaranha 1
- Victor Herbst 1
- Peter Ljunglöf 1
- Ana Carolina Luz 1
- Ana Carolina Loçasso Luz 1
- Ben Lyngfelt 1
- Yulla Liquer Navarro 1
- Juliana de Oliveira 1
- Márcio Henrique Pelegrino de Freitas 1
- Luiz Fernando Pereira 1
- Sofia Reinach 1
- Erik Santos 1
- Natalia Sigiliano 1
- Natália Sathler Sigiliano 1
- Mariana Souza 1
- Mariana Souza Mota 1
- Peter Uhrig 1
- Janina Wildfeuer 1
- Alexander Ziem 1