Tolga MAVİGÖZ, Adem ALTIN, Merve YENİÇERİ ÖZATA
Journal of Dental Sciences and Education - 2026;4(1):17-22
Aims: The aim of this study is to evaluate the correctness rate/probability of answers provided by AI-powered chatbots (Gemini Advanced 2.5 Pro, ChatGPT-4 omni, ChatGPT-5 and DeepSeek v3) to single-answer, multiple-choice basic sciences questions from the Dentistry Specialization Examination (DUS) administered between 2012 and 2025. Methods: A total of 539 multiple-choice questions from the basic sciences section of the DUS from 2012 to 2025 were used. Each question was presented directly in a new session. The rates of correct/incorrect answers were calculated based on the subject of the question, the year it was asked, and the chatbot model. The rate and probability of incorrect answers were evaluated using chi-square and binary logistic regression analyses. Results: The rate of incorrect answers was highest in 2012, with a significant decrease observed in subsequent years (p<0.05). Among the subjects, Anatomy had the highest rate of incorrect answers, while Pathology had the lowest (p<0.05). When comparing chatbot models, Gemini Advanced 2.5 Pro was found to have a statistically significantly lower error rate than ChatGPT-4 omni and DeepSeek v3 (p<0.05). In the regression analysis, the risk of providing an incorrect answer was statistically significantly higher for the years 2012 and 2018; for the fields of Anatomy, Physiology and Microbiology; and for the ChatGPT-4 omni and DeepSeek v3 models, compared to their respective reference groups (p<0.05). Conclusion: AI-powered chatbots provided more accurate answers to more recent questions. In the subject-based performance analysis, the rate of incorrect answers for Anatomy questions was high. In the chatbot comparison, Gemini Advanced 2.5 Pro produced more accurate answers than DeepSeek v3. While AI-powered chatbots can be a potential supplementary tool in the preparation process for DUS, their accuracy and competence across different subjects are limited.