CONSULTATION POTENTIAL OF ARTIFICIAL INTELLIGENCE CHATBOTS IN PROSTATITIS MANAGEMENT: AN EVALUATION OF QUALITY, RELIABILITY AND READABILITY

Halil DEMİRÇAKAN, Burak KÖSEOĞLU, Taha Numan YIKILMAZ

Journal of Urological Surgery - 2026;13(3):178-185

Çumra State Hospital, Clinic of Urology, Konya, Türkiye

 

Objective: This study aimed to assess the responses of three artificial intelligence (AI) chatbots (ChatGPT-4, Gemini Pro, Llama 3.1 Large) on prostatitis using quality, reliability, readability scales and examine their role in disease management. Materials and Methods: Keywords related to prostatitis were identified using Google Trends and Semrush platforms. The search volume and regional distribution of these terms over the past five years were analyzed, leading to the selection of 25 questions. The core question set was categorized into six subgroups: general information, symptoms, diagnostic methods, treatment methods, complications, and myths. The responses generated by the AI chatbots were assessed for readability using the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease (FRES) scores. Quality and reliability were evaluated using the Educational Quality of Information for Patients (EQIP) score and the Modified DISCERN score. Results: No significant difference was observed among AI chatbots in mean FKGL scores (p=0.354). However, ChatGPT-4 had a significantly higher mean FRES score than Gemini Pro and Llama 3.1 Large (p=0.016 and p=0.003, respectively). Gemini Pro had the highest mean EQIP score, significantly surpassing Llama 3.1 Large and ChatGPT-4 (p<0.001); Llama 3.1 Large had the highest median Modified DISCERN score (p<0.001). Across all subgroup analyses, Gemini Pro yielded the highest mean EQIP score, while Llama 3.1 Large had the highest median Modified DISCERN score. Conclusion: The findings of this study suggest that Llama 3.1 Large provides more reliable responses to questions about prostatitis, whereas Gemini Pro delivers higher-quality responses. However, the readability levels of all AI chatbots exceeded the recommended 6th-grade level, indicating that their responses may be challenging for general audiences.