Logo image
ChatGPT Versus Custom-Trained Chatbot for Urogynecology Surgery Counseling
Journal article   Peer reviewed

ChatGPT Versus Custom-Trained Chatbot for Urogynecology Surgery Counseling

Leanne Brechtel, Colin Johnson, Sarah Rabice, Douglas Russo, Kimberly A Kenne, Christina Lewicky-Gaupp and Joseph T Kowalski
Urogynecology
05/25/2026
DOI: 10.1097/SPV.0000000000001862
PMID: 42240077

View Online

Abstract

Patient education materials are important components of shared decision making in urogynecologic surgery. Traditional materials are often difficult to update, lack personalization, and are not easily adaptable. Chatbots offer a new approach to generating comprehensive and accessible content, but their utility in a clinic setting remains unclear. The objective of this study was to compare the performance of a general-purpose chatbot, ChatGPT, and a domain-specific chatbot developed by the Foundation for Female Health Awareness (FFHA) (FFHA Assistant) in generating surgical counseling information, using standardized materials from the International Urogynecological Association (IUGA) as reference. Seven IUGA handouts representing common urogynecologic surgical procedures were selected. Identical prompts were submitted to ChatGPT-4.0 and the FFHA Assistant. Responses were reviewed by 7 blinded urogynecology experts using 5-point Likert scales to assess accuracy, completeness, and understandability. Readability was evaluated using the Flesch-Kincaid Grade Level and Flesch Reading Ease Score. ChatGPT-4.0 outperformed in completeness as compared with the IUGA leaflets (median 4 [3-5] vs 3 [3-4], P<0.01), whereas the FFHA Assistant scored higher in accuracy (median 3 [3-3] vs 3 [2-3], P<0.01) and understandability (3 [3-4] vs 3 [3-3], P<0.01). Both large language models generated longer responses than the IUGA leaflets. The FFHA Assistant responses had better readability scores, aligning more closely with health literacy recommendations. Both chatbots generated counseling content comparable to or superior to existing materials. The domain-specific FFHA Assistant responses were better aligned with health literacy recommendations. Further research is needed to better understand the reproducibility of responses and the clinical utility of chatbots in patient education.

Details

Metrics

1 Record Views
Logo image