Sirui Chen
Author directoryPapers on this page may belong to the following people: Sirui Chen, Sirui Chen
2025
PreGenie: An Agentic Framework for High-quality Visual Presentation Generation
Xiaojie Xu | Xinli Xu | Sirui Chen | Haoyu Chen | Fan Zhang | Ying-Cong Chen
Findings of the Association for Computational Linguistics: EMNLP 2025
Xiaojie Xu | Xinli Xu | Sirui Chen | Haoyu Chen | Fan Zhang | Ying-Cong Chen
Findings of the Association for Computational Linguistics: EMNLP 2025
Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations restrict their application in formal contexts like business and scientific research. To address these challenges, we propose PreGenie, an agentic and modular framework powered by multimodal large language models (MLLMs) for generating high-quality visual presentations.PreGenie is built on the Slidev presentation framework, where slides are rendered from Markdown code. It operates in two stages: (1) Analysis and Initial Generation, which summarizes multimodal input and generates initial code, and (2) Review and Re-generation, which iteratively reviews intermediate code and rendered slides to produce final, high-quality presentations. Each stage leverages multiple MLLMs that collaborate and share information. Comprehensive experiments demonstrate that PreGenie excels in multimodal understanding, outperforming existing models in both aesthetics and content consistency, while aligning more closely with human design preferences.
2024
CLEAR: Can Language Models Really Understand Causal Graphs?
Sirui Chen | Mengying Xu | Kun Wang | Xingyu Zeng | Rui Zhao | Shengjie Zhao | Chaochao Lu
Findings of the Association for Computational Linguistics: EMNLP 2024
Sirui Chen | Mengying Xu | Kun Wang | Xingyu Zeng | Rui Zhao | Shengjie Zhao | Chaochao Lu
Findings of the Association for Computational Linguistics: EMNLP 2024
Causal reasoning is a cornerstone of how humans interpret the world. To model and reason about causality, causal graphs offer a concise yet effective solution. Given the impressive advancements in language models, a crucial question arises: can they really understand causal graphs? To this end, we pioneer an investigation into language models’ understanding of causal graphs. Specifically, we develop a framework to define causal graph understanding, by assessing language models’ behaviors through four practical criteria derived from diverse disciplines (e.g., philosophy and psychology). We then develop CLEAR, a novel benchmark that defines three complexity levels and encompasses 20 causal graph-based tasks across these levels. Finally, based on our framework and benchmark, we conduct extensive experiments on six leading language models and summarize five empirical findings. Our results indicate that while language models demonstrate a preliminary understanding of causal graphs, significant potential for improvement remains.