Hasan ÖNNER, Lütfü PERKTAŞ, Kevser ÖKSÜZOĞLU, Farise YILMAZ, Fettah EREN, Şerefnur ÖZTÜRK
Van Medical Journal - 2026;33(3):277-283
Introduction: To evaluate the diagnostic performance of large language models (LLMs) with FDG PET-derived Z-score profiles and structured clinical information in patients with suspected neurodegenerative diseases (NDs). Materials and Methods: This retrospective study included patients who underwent FDG-PET imaging of the brain for suspected neurodegenerative diseases. FDG PET brain imaging Z-score values derived from a database and anonymized structured clinical information were provided to four LLMs (ChatGPT, Grok, Gemini, and DeepSeek). Each model generated a single diagnosis among Alzheimer's disease, frontotemporal dementia, dementia with Lewy bodies, vascular dementia, primary progressive aphasia, or normal/nonspecific. A multidisciplinary consensus diagnosis served as the reference standard. LLM outputs were stratified into subgroups: overall, high diagnostic confidence (HC; >=85%), epicenter concordance (EC), and combined (HC + EC). Diagnostic agreement was assessed using Cohen's kappa (kappa). Results: A total of 80 patients (42 females, 38 males) were included, with multidisciplinary consensus diagnoses of Alzheimer's disease (n=28), frontotemporal dementia (n=17), normal/nonspecific findings (n=17), dementia with Lewy bodies (n=8), primary progressive aphasia (n=8), and vascular dementia (n=2). All LLMs showed statistically significant agreement with the multidisciplinary consensus diagnosis. Across the overall cohort, ChatGPT had the highest agreement (kappa=0.760), followed by Grok (kappa=0.648) and DeepSeek (kappa=0.639). In the HC subgroup, agreement improved across all models, with ChatGPT reaching kappa=0.872, followed by DeepSeek (kappa=0.789), Grok (kappa=0.711), and Gemini (kappa=0.676). In the EC subgroup, agreement remained high, with ChatGPT (kappa=0.828) and Grok (kappa=0.799) showing substantial concordance, and DeepSeek (kappa=0.682) and Gemini (kappa=0.542) demonstrating moderate agreement. In the combined HC + EC subgroup, ChatGPT achieved the strongest performance (kappa=0.860), followed by Grok (kappa=0.824) and DeepSeek (kappa=0.786), while Gemini reached moderate agreement (kappa=0.652). Conclusion: LLMs showed moderate-to-high agreement with multidisciplinary consensus diagnoses in patients with suspected neurodegenerative diseases. Agreement was higher in cases with high diagnostic confidence and epicenter concordance. These findings suggest that LLMs may have potential as adjunctive decision-support tools in FDG PET-based neuroimaging for NDs.