Sailaja Rajanala
2021
Auditing Keyword Queries Over Text Documents
Bharath Kumar Reddy Apparreddy
|
Sailaja Rajanala
|
Manish Singh
Proceedings of the 18th International Conference on Natural Language Processing (ICON)
Data security and privacy is an issue of growing importance in the healthcare domain. In this paper, we present an auditing system to detect privacy violations for unstructured text documents such as healthcare records. Given a sensitive document, we present an anomaly detection algorithm that can find the top-k suspicious keyword queries that may have accessed the sensitive document. Since unstructured healthcare data, such as medical reports and query logs, are not easily available for public research, in this paper, we show how one can use the publicly available DBLP data to create an equivalent healthcare data and query log, which can then be used for experimental evaluation.