PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents

Minjia Wang; Yunfeng Wang; Xiao Ma; Dexin Lv; Qifan Guo; Lynn Zheng; Benliang Wang; Lei Wang; Jiannan Li; Yongwei Xing; Junzhe Xu; Zheng Sun

PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents

Minjia Wang, Yunfeng Wang, Xiao Ma, Dexin Lv, Qifan Guo, Lynn Zheng, Benliang Wang, Lei Wang, Jiannan Li, Yongwei Xing, Junzhe Xu, Zheng Sun

Abstract

Digital footprints—records of individuals’ interactions with digital systems—are essential for studying behavior, developing personalized applications, and training machine learning models. However, research in this area is often hindered by the scarcity of diverse and accessible data. To address this limitation, we propose a novel method for synthesizing realistic digital footprints using large language model (LLM) agents. Starting from a structured user profile, our approach generates diverse and plausible sequences of user events, ultimately producing corresponding digital artifacts such as emails, messages, calendar entries, reminders, etc. Intrinsic evaluation results demonstrate that the generated dataset is more diverse and realistic than existing baselines. Moreover, models fine-tuned on our synthetic data outperform those trained on other synthetic datasets when evaluated on real-world out-of-distribution tasks.

Anthology ID:: 2026.eacl-industry.5
Volume:: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track)
Month:: March
Year:: 2026
Address:: Rabat, Morocco
Editors:: Yevgen Matusevych, Gülşen Eryiğit, Nikolaos Aletras
Venue:: EACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 60–77
Language:
URL:: https://aclanthology.org/2026.eacl-industry.5/
DOI:
Bibkey:
Cite (ACL):: Minjia Wang, Yunfeng Wang, Xiao Ma, Dexin Lv, Qifan Guo, Lynn Zheng, Benliang Wang, Lei Wang, Jiannan Li, Yongwei Xing, Junzhe Xu, and Zheng Sun. 2026. PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 60–77, Rabat, Morocco. Association for Computational Linguistics.
Cite (Informal):: PersonaTrace: Synthesizing Realistic Digital Footprints with LLM Agents (Wang et al., EACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.eacl-industry.5.pdf

PDF Cite Search Fix data