Ahmet Kürşat KARAMAN, Naz Günay DEMİRCAN, Berkay YILMAZ, Tufan Agah KARTUM, Ahu Senem DEMİRÖZ, Ayşe Mine ÖNENERK, Bora KORKMAZER, Osman KIZILKILIÇ
Cerrahpaşa Medical Journal - 2026;50(1):1-9
Objective: To evaluate the diagnostic performance of 2 multimodal large language models (LLMs), ChatGPT-4V and Claude Sonnet 3.7, in preoperative magnetic resonance imaging (MRI)-based diagnosis and histopathological subtyping of WHO Grade I meningiomas and to compare their performance with a radiology resident. Methods: This retrospective study included 113 patients with WHO Grade I meningiomas who underwent preoperative MRI between January 2020 and December 2024. Multisequence MRI (T1WI, T2WI, FLAIR, T1WI+C, DWI, ADC maps) and radiologist-extracted imaging features were provided to both LLMs in a structured "images + human extracted features" format. Each evaluator performed tumor type and subtype prediction using the 2021 WHO CNS classification. Diagnostic accuracies with 95% confidence intervals were calculated, and comparisons between LLMs and the resident were performed using McNemar's test. Results: For type classification, ChatGPT-4V achieved 76.99% accuracy and Claude Sonnet 3.7 achieved 73.45%, while the radiology resident showed significantly higher accuracy at 92.04% (P < .001 for both). Subtype classification was challenging. ChatGPT-4V achieved 20.35% accuracy, Claude Sonnet 3.7 30.09%, and the resident 30.97%, with no statistically significant differences among them. No significant performance differences were observed between the 2 LLMs. Conclusion: Multimodal LLMs demonstrated moderate ability to identify WHO Grade I meningiomas but did not match the diagnostic performance of the radiology resident. Subtype prediction remained limited, reflecting the constraints of conventional MRI for distinguishing subtle histopathological subtypes. Large language models may assist preoperative assessment, but expert interpretation remains essential. Further development using larger datasets and domain-specific model refinement may improve clinical usefulness.