Abstract
Natural Language Processing (NLP)-based data extraction from electronic health records (EHRs) holds significant potential to simplify clinical management and aid research. This review aims to evaluate the current landscape of NLP-based data extraction in prostate cancer (PCa) management.
We conducted a literature search of PubMed and Google Scholar databases using the keywords: "Natural Language Processing," "Prostate Cancer," "data extraction," and "EHR." with variations of each. No language or time limits were imposed. All results were collected in a standardized fashion, including country of origin, sample size, algorithm, objective of outcome, and model performance. The precision, Recall, and the F1 score of studies were collected as a metric of model performance.
Of 14 studies included in the review, two articles focused on documenting digital rectal exams, one on identifying and quantifying pain secondary to PCa, eight on extracting staging/grading information from clinical reports, with an emphasis on TNM-classification, risk stratification, and identifying metastasis, two articles focused on patient-centered post-treatment outcomes like incontinence, erectile and bowel-dysfunction, and one on loneliness/social isolation following PCa diagnosis. All models showed moderate to high data annotation/extraction accuracy compared to the gold standard method of manual data extraction by chart review. Despite their potential, NLPs face challenges in handling ambiguous, institution-specific language and context nuances, leading to occasional inaccuracies in clinical data interpretation.
NLP-based data extraction has successfully extracted various outcomes from PCa patients' EHRs. It holds the potential for automating outcome monitoring and data collection, resulting in time and labor savings.