Contrastively Pre-trained Event Embeddings with Schema-free LLM Annotations

Frank Mtumbuka, Steven Schockaert


Abstract
Event extraction is a notoriously challenging problem, among others due to the scarcity of suitable training data. Moreover, event-centric knowledge bases are not available for most domains, making traditional distant supervision strategies difficult to implement. In this paper, we evaluate the potential of using LLM-generated annotations as an alternative distant supervision signal. Specifically, we create a synthetically labelled event extraction corpus, using an LLM to identify event triggers and arguments, and to provide corresponding free-text descriptions. We then pre-train event embedding models on this corpus using a contrastive loss, before fine-tuning them in the usual way. We empirically show the effectiveness of this approach.
Anthology ID:
2026.lrec-1.591
Volume:
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Month:
May
Year:
2026
Address:
Palma de Mallorca, Spain
Editors:
Stelios Piperidis, Núria Bel, Henk van den Heuvel, Nancy Ide, Simon Krek, Antonio Toral
Venue:
LREC
SIG:
Publisher:
ELRA Language Resource Association
Note:
Pages:
7457–7478
Language:
External URL:
https://lrec.elra.info/lrec2026-main-591
DOI:
10.63317/3sezhi63dcqv
Bibkey:
Cite (ACL):
Frank Mtumbuka and Steven Schockaert. 2026. Contrastively Pre-trained Event Embeddings with Schema-free LLM Annotations. In Proceedings of the Fifteenth Language Resources and Evaluation Conference, pages 7457–7478, Palma de Mallorca, Spain. ELRA Language Resource Association.
Cite (Informal):
Contrastively Pre-trained Event Embeddings with Schema-free LLM Annotations (Mtumbuka & Schockaert, LREC 2026)
Copy Citation: