Tuğçe PAKSOY, Selin GAŞ, Seval CEYLAN ŞEN, Özlem SARAÇ ATAGÜN, Gülbahar USTAOĞLU, Şeyma ÇARDAKÇI BAHAR, Savaş ÖZARSLANTÜRK
Journal of Dental Sciences and Education - 2026;4(3):76-83
Aims: This study aimed to identify the clinically relevant patient questions about apical root-end resection based on international consensus guidelines, and to systematically evaluate the accuracy, quality, usefulness, and readability of generated responses by four Artificial Intelligence (AI)-powered conversational agents, ChatGPT-4, DeepSeek V3.1, Gemini 3, and Copilot. Methods: A standardized set of 20 clinical questions was developed from the American Association of Endodontists guidelines and international consensus reports. A priori, an evidence-based answer key was established as the reference standard. Queries were submitted to ChatGPT-4, DeepSeek V3.1, Gemini 3, and Copilot, following a standardized interaction protocol designed to minimize personalization and carryover effects, such as incognito mode, newly created accounts, and no response regeneration. The outputs were independently assessed by three experts in periodontics and oral radiology, using validated tools, such as CLEAR criteria, modified Global Quality Score (mGQS), accuracy ratings, DISCERN, readability indices (Flesch reading ease, FRE and Flesch-Kincaid grade level, FKGL) and the patient education materials assessment tool (PEMAT). Results: ChatGPT-4 and Copilot achieved significantly higher CLEAR scores than Gemini 3. Copilot achieved significantly higher mGQS and accuracy than DeepSeek V3.1 and Gemini 3 and received significantly lower, and therefore more favorable, usefulness scores. Gemini 3 scored higher on the FKGL test and lower on the understandability test compared to DeepSeek V3.1. Relationships between evaluation metrics were different between models. Conclusion: The AI models evaluated demonstrated significant variability in clinical accuracy, quality, readability, and patient-centered usability. Copilot had the best overall mGQS and Accuracy scores and the best Usefulness scores, while the models varied in their strengths in readability, understandability, and actionability.