Sevim TÜRAY, Nefise ARIBAŞ ÖZ, Elif Meliha SÖZBİR, Veysel ULUDAĞ
Experimental Biomedical Research - 2026;9(3):202-211
Aim: Artificial intelligence (AI) and large language models (LLMs) are increasingly used in medicine, yet their role in pediatric neurology remains unclear. Given the unique challenges in this field, evaluating the accuracy, reliability, and clinical applicability of AI-generated responses is essential. This study aimed to compare the performance of ChatGPT-3.5 and ChatGPT-4.0 in clinical decision-making in pediatric neurology based on expert evaluations using standardized rating criteria. Methods: This prospective observational study included 61 board-certified pediatric neurologists who assessed AI-generated responses to ten common pediatric neurology cases. Responses were evaluated on a five-point Likert scale for accuracy, reliability, clinical applicability, and comprehensibility. Since the same raters evaluated both models, comparisons were performed using the non-parametric Wilcoxon signed-rank test. Internal consistency was examined with Cronbach's alpha, and correlations between experience and model ratings were analyzed using Pearson's test. Results: GPT-4.0 achieved higher mean scores than GPT-3.5 across all domains, particularly for accuracy (4.2 +/- 0.4 vs 3.5 +/- 0.5, p = 0.38) and comprehensibility (4.3 +/- 0.4 vs 3.6 +/- 0.5, p = 0.62). Although GPT-4.0 performed slightly better overall, none of the differences were statistically significant (p > 0.05). Cronbach's alpha values indicated low internal consistency (ranging from 0.58 to 0.67), suggesting variability among raters. Expert experience showed no significant correlation with ratings, implying that AI evaluations were largely experienced-independent. Conclusion: While GPT-4.0 demonstrated modest improvements over GPT-3.5, neither model achieved sufficient reliability for independent clinical use in pediatric neurology. Addressing AI hallucinations, enhancing internal consistency in expert evaluations, and promoting safe clinical integration remain key priorities for future research.