Peng Wu
Also published as: Wupeng Njust
2025
Exploring and Detecting Self-disclosure in Multi-modal posts on Chinese Social Media
Jingbao Luo | Ming Liu | Aoli Huo | Fujing Hu | Gang Li | Wupeng Njust
Findings of the Association for Computational Linguistics: EMNLP 2025
Jingbao Luo | Ming Liu | Aoli Huo | Fujing Hu | Gang Li | Wupeng Njust
Findings of the Association for Computational Linguistics: EMNLP 2025
Self-disclosure can provide psychological comfort and social support, but it also carries the risk of unintentionally revealing sensitive information, leading to serious privacy concerns. Research on self-disclosure in Chinese multimodal contexts remains limited, lacking high-quality corpora, analysis, and methods for detection. This work focuses on self-disclosure behaviors on Chinese multimodal social media platforms and constructs a high-quality text-image corpus to address this critical data gap. We systematically analyze the distribution of self-disclosure types, modality preferences, and their relationship with user intent, uncovering expressive patterns unique to the Chinese multimodal context. We also fine-tune five multimodal large language models to enhance self-disclosure detection in multimodal scenarios. Among these models, the Qwen2.5-omni-7B achieved a strong performance, with a partial span F1 score of 88.2%. This study provides a novel research perspective on multimodal self-disclosure in the Chinese context.
Can Language Models Capture Human Writing Preferences for Domain-Specific Text Summarization?
Jingbao Luo | Ming Liu | Ran Liu | Yongpan Sheng | Xin Hu | Gang Li | Peng Wu
Findings of the Association for Computational Linguistics: ACL 2025
Jingbao Luo | Ming Liu | Ran Liu | Yongpan Sheng | Xin Hu | Gang Li | Peng Wu
Findings of the Association for Computational Linguistics: ACL 2025
With the popularity of large language models and their high-quality text generation capabilities, researchers are using them as auxiliary tools for text summary writing. Although summaries generated by these large language models are smooth and capture key information sufficiently, the quality of their output depends on the prompt, and the generated text is somewhat procedural to a certain extent. We construct LecSumm to verify whether language models truly capture human writing preferences, in which we recruit 200 college students to write summaries for lecture notes on ten different machine-learning topics and analyze writing preferences in real-world human summaries through the dimensions of length, content depth, tone & style, and summary format. We define the method of capturing human writing preferences by language models as finetuning pre-trained models with data and designing prompts to optimize the output of large language models. The results of translating the analyzed human writing preferences into prompts and conducting experiments show that both models still fail to capture human writing preferences effectively. Our LecSumm dataset brings new challenges to finetuned and prompt-based large language models on the task of human-centered text summarization.
2020
GCDST: A Graph-based and Copy-augmented Multi-domain Dialogue State Tracking
Peng Wu | Bowei Zou | Ridong Jiang | AiTi Aw
Findings of the Association for Computational Linguistics: EMNLP 2020
Peng Wu | Bowei Zou | Ridong Jiang | AiTi Aw
Findings of the Association for Computational Linguistics: EMNLP 2020
As an essential component of task-oriented dialogue systems, Dialogue State Tracking (DST) takes charge of estimating user intentions and requests in dialogue contexts and extracting substantial goals (states) from user utterances to help the downstream modules to determine the next actions of dialogue systems. For practical usages, a major challenge to constructing a robust DST model is to process a conversation with multi-domain states. However, most existing approaches trained DST on a single domain independently, ignoring the information across domains. To tackle the multi-domain DST task, we first construct a dialogue state graph to transfer structured features among related domain-slot pairs across domains. Then, we encode the graph information of dialogue states by graph convolutional networks and utilize a hard copy mechanism to directly copy historical states from the previous conversation. Experimental results show that our model improves the performances of the multi-domain DST baseline (TRADE) with the absolute joint accuracy of 2.0% and 1.0% on the MultiWOZ 2.0 and 2.1 dialogue datasets, respectively.
2019
Learning Representation Mapping for Relation Detection in Knowledge Base Question Answering
Peng Wu | Shujian Huang | Rongxiang Weng | Zaixiang Zheng | Jianbing Zhang | Xiaohui Yan | Jiajun Chen
Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Peng Wu | Shujian Huang | Rongxiang Weng | Zaixiang Zheng | Jianbing Zhang | Xiaohui Yan | Jiajun Chen
Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Relation detection is a core step in many natural language process applications including knowledge base question answering. Previous efforts show that single-fact questions could be answered with high accuracy. However, one critical problem is that current approaches only get high accuracy for questions whose relations have been seen in the training data. But for unseen relations, the performance will drop rapidly. The main reason for this problem is that the representations for unseen relations are missing. In this paper, we propose a simple mapping method, named representation adapter, to learn the representation mapping for both seen and unseen relations based on previously learned relation embedding. We employ the adversarial objective and the reconstruction objective to improve the mapping performance. We re-organize the popular SimpleQuestion dataset to reveal and evaluate the problem of detecting unseen relations. Experiments show that our method can greatly improve the performance of unseen relations while the performance for those seen part is kept comparable to the state-of-the-art.
2011
Improving the Accessibility of Line Graphs in Multimodal Documents
Charles F. Greenbacker | Peng Wu | Sandra Carberry | Kathleen F. McCoy | Stephanie Elzer | David D. McDonald | Daniel Chester | Seniz Demir
Proceedings of the Second Workshop on Speech and Language Processing for Assistive Technologies
Charles F. Greenbacker | Peng Wu | Sandra Carberry | Kathleen F. McCoy | Stephanie Elzer | David D. McDonald | Daniel Chester | Seniz Demir
Proceedings of the Second Workshop on Speech and Language Processing for Assistive Technologies
Search
Fix author
Co-authors
- Sandra Carberry 2
- Stephanie Elzer 2
- Gang Li 2
- Ming Liu 2
- Jingbao Luo 2
- Kathleen F. McCoy 2
- Aiti Aw 1
- Jiajun Chen 1
- Daniel Chester 1
- Seniz Demir 1
- Charles Greenbacker 1
- Charles F. Greenbacker 1
- Fujing Hu 1
- Xin Hu 1
- Shujian Huang (书剑 黄) 1
- Aoli Huo 1
- Ridong Jiang 1
- Ran Liu 1
- David D. McDonald 1
- Yongpan Sheng 1
- Rongxiang Weng 1
- Xiaohui Yan 1
- Jianbing Zhang 1
- Zaixiang Zheng 1
- Bowei Zou (邹博伟) 1