Efficient Encoding of Pathology Reports Using Natural Language Processing

Rebecka Weegar, Jan F Nygård, Hercules Dalianis


Abstract
In this article we present a system that extracts information from pathology reports. The reports are written in Norwegian and contain free text describing prostate biopsies. Currently, these reports are manually coded for research and statistical purposes by trained experts at the Cancer Registry of Norway where the coders extract values for a set of predefined fields that are specific for prostate cancer. The presented system is rule based and achieves an average F-score of 0.91 for the fields Gleason grade, Gleason score, the number of biopsies that contain tumor tissue, and the orientation of the biopsies. The system also identifies reports that contain ambiguity or other content that should be reviewed by an expert. The system shows potential to encode the reports considerably faster, with less resources, and similar high quality to the manual encoding.
Anthology ID:
R17-1100
Volume:
Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017
Month:
September
Year:
2017
Address:
Varna, Bulgaria
Editors:
Ruslan Mitkov, Galia Angelova
Venue:
RANLP
SIG:
Publisher:
INCOMA Ltd.
Note:
Pages:
778–783
Language:
URL:
https://doi.org/10.26615/978-954-452-049-6_100
DOI:
10.26615/978-954-452-049-6_100
Bibkey:
Cite (ACL):
Rebecka Weegar, Jan F Nygård, and Hercules Dalianis. 2017. Efficient Encoding of Pathology Reports Using Natural Language Processing. In Proceedings of the International Conference Recent Advances in Natural Language Processing, RANLP 2017, pages 778–783, Varna, Bulgaria. INCOMA Ltd..
Cite (Informal):
Efficient Encoding of Pathology Reports Using Natural Language Processing (Weegar et al., RANLP 2017)
Copy Citation:
PDF:
https://doi.org/10.26615/978-954-452-049-6_100