PII-Bench: Evaluating Query-Aware Privacy Protection Systems

Hao Shen; Zhouhong Gu; Haokai Hong; Weili Han; Hongfeng Chai (柴洪峰)

PII-Bench: Evaluating Query-Aware Privacy Protection Systems

Hao Shen, Zhouhong Gu, Haokai Hong, Weili Han, Hongfeng Chai

Abstract

The widespread adoption of Large Language Models (LLMs) has raised significant privacy concerns regarding the exposure of personally identifiable information (PII) in user prompts. To address this challenge, we propose a query-unrelated PII masking strategy and introduce PII-Bench, the first comprehensive evaluation framework for assessing privacy protection systems. PII-Bench comprises 2,842 test samples across 7 PII types with 55 fine-grained subcategories, featuring diverse scenarios from single-subject descriptions to complex multi-party interactions. Each sample is carefully crafted with a user query, context description, and standard answer indicating query-relevant PII. Our empirical evaluation reveals that while current models perform adequately in basic PII detection, they show significant limitations in determining PII query relevance. Even advanced LLMs struggle with this task, particularly in handling complex multi-subject scenarios, indicating substantial room for improvement in achieving intelligent PII masking.

Anthology ID:: 2026.acl-long.227
Volume:: Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
Month:: July
Year:: 2026
Address:: San Diego, California, United States
Editors:: Maria Liakata, Viviane P. Moreira, Jiajun Zhang, David Jurgens
Venue:: ACL
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 4991–5026
Language:
URL:: https://aclanthology.org/2026.acl-long.227/
DOI:
Bibkey:
Cite (ACL):: Hao Shen, Zhouhong Gu, Haokai Hong, Weili Han, and Hongfeng Chai. 2026. PII-Bench: Evaluating Query-Aware Privacy Protection Systems. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4991–5026, San Diego, California, United States. Association for Computational Linguistics.
Cite (Informal):: PII-Bench: Evaluating Query-Aware Privacy Protection Systems (Shen et al., ACL 2026)
Copy Citation:
PDF:: https://aclanthology.org/2026.acl-long.227.pdf
Checklist:: 2026.acl-long.227.checklist.pdf

PDF Cite Search Checklist Fix data