Bence Sárossy
Also published as: Bence Sarossy
2026
Managing Growth in a National Corpus: The Hungarian National Corpus 3.0 (MNSZ3)
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
Noémi Ligeti-Nagy | Enikő Héja | Ágnes Bánfi | Flóra Földesi | Bence Sárossy | Boglárka Skrabák | Tamás Váradi | Gábor Prószéky
Proceedings of the 12th Workshop on Challenges in the Management of Large Corpora
The third generation of the Hungarian National Corpus (MNSZ3) aims to provide a large-scale, curated, and well-described corpus resource needed for the sustainable digital presence of Hungarian. Building on the domain structure and proportions of MNSZ2 (v2.0.5; 1.04 billion running words), the project targets a substantial increase in scale while also strengthening the coverage and metadata description of Hungarian language use outside Hungary. MNSZ3 retains the six traditional domains of the earlier corpus—press, fiction, scientific, official, personal, and transcribed spoken language—and is planned to reach approximately 10 billion tokens. This paper presents the motivation and design principles of the project, outlines the practical decisions and procedures used in data collection and cleaning, and discusses the annotation strategy developed for large-scale processing. In planning the linguistic analysis, we build on the complementary strengths of HuSpaCy and e-magyar: HuSpaCy provides the unified and efficient UD-oriented processing backbone, while e-magyar (emMorph) is preserved as an explicit additional layer for morphology and lemmatisation.
2025
HuGME: A benchmark system for evaluating Hungarian generative LLMs
Noémi Ligeti-Nagy | Gabor Madarasz | Flora Foldesi | Mariann Lengyel | Matyas Osvath | Bence Sarossy | Kristof Varga | Győző Zijian Yang | Enikő Héja | Tamás Váradi | Gábor Prószéky
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
Noémi Ligeti-Nagy | Gabor Madarasz | Flora Foldesi | Mariann Lengyel | Matyas Osvath | Bence Sarossy | Kristof Varga | Győző Zijian Yang | Enikő Héja | Tamás Váradi | Gábor Prószéky
Proceedings of the Fourth Workshop on Generation, Evaluation and Metrics (GEM²)
In this study, we introduce the Hungarian Generative Model Evaluation (HuGME) benchmark, a new framework designed to assess the linguistic proficiency of large language models (LLMs) in Hungarian. HuGME evaluates models across a diverse set of linguistic and reasoning skills, including bias, toxicity, faithfulness, relevance, summarization, prompt alignment, readability, spelling, grammaticality, and domain-specific knowledge through tasks like TruthfulQA and MMLU. We applied HuGME to a range of Hungarian LLMs, including those developed in-house as well as several publicly available models that claim Hungarian language proficiency. This paper presents the comparative results of these evaluations, shedding light on the capabilities of current LLMs in processing the Hungarian language. Through our analysis, we aim to both showcase the current state of Hungarian linguistic processing in LLMs and provide a foundational resource for future advancements in the field.