Insight discovery in structured data

Kristyna Onderkova


Abstract
My research focuses on improving textual inference in large language models (LLMs) for natural language generation, particularly in data-to-text generation. While LLMs are increasingly used to generate reports and insights from data, they often produce factually inaccurate or shallow outputs, limiting their usefulness. I work on integrating LLMs with symbolic operations through code generation for deeper and more faithful inferences. As generation tasks are often under-specified, both models and humans rely on implicit presuppositions, and mismatches can lead to errors or misinterpretation. I investigate how such presuppositions affect generation outputs and evaluation, how human presuppositions shape the perceived interestingness of the insights, and how they can be leveraged to improve insight generation.
Anthology ID:
2025.ynlg-main.4
Volume:
Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation
Month:
October
Year:
2025
Address:
Hanoi, Vietnam
Editors:
Alyssa Allen, Nils Feldhus, Rudali Huidrom, Michela Lorandi, Adarsa Sivaprasad, Patrícia Schmidtová
Venue:
YNLG
SIG:
SIGGEN
Publisher:
Association for Computational Linguistics
Note:
Pages:
17–20
Language:
URL:
https://aclanthology.org/2025.ynlg-main.4/
DOI:
Bibkey:
Cite (ACL):
Kristyna Onderkova. 2025. Insight discovery in structured data. In Proceedings of the 1st Workshop for Young Researchers in Natural Language Generation, pages 17–20, Hanoi, Vietnam. Association for Computational Linguistics.
Cite (Informal):
Insight discovery in structured data (Onderkova, YNLG 2025)
Copy Citation:
PDF:
https://aclanthology.org/2025.ynlg-main.4.pdf