Fatma HAKYEMEZ, Emine CIHAN, Cansu Sahbaz PIRINCCI
Eskisehir Medical Journal - 2026;7(3):392-396
Introduction: This research aimed to comparatively evaluate the performance of ChatGPT and Gemini in terms of medical accuracy and readability regarding first aid in musculoskeletal system traumas. Methods: In the research, the responses given by both artificial intelligence systems to 10 first aid questions directed in the same format and in the Turkish language were analyzed. The responses were evaluated by 7 experts with clinical experience in the fields of first aid training and emergency health services. Flesch-Kincaid Grade Level scores were recorded for each response. Mann-Whitney U test was used in statistical analyses, and the significance level was accepted as p<0.05. Results: In questions 3, 4, and 6, Gemini's response quality was found to be statistically significantly superior (p=0.035, p=0.035, and p=0.037). In the readability analysis, Flesch-Kincaid scores ranged between 9.85-13.46 for ChatGPT and 12.28-17.35 for Gemini. It was determined that ChatGPT responses were more readable, while Gemini responses required a higher reading level. The findings indicate an inverse relationship between the two models in terms of accuracy and readability. Conclusion: Artificial intelligence-powered chatbots hold significant public health potential for rapid access to first aid information. However, it has been observed that more accurate responses are not always more readable. Therefore, it is thought that optimizing both the accuracy and readability of artificial intelligence-based health information will contribute to its safe and effective use by the public.