Mehmet YILDIRIM, Tülay Dilara DEMİRAY, Arda AYTEN, Esat Kıvanç KAYA
Sakarya Tıp Dergisi - 2026;16(2):200-207
Objective: Family members of patients admitted to the intensive care unit (ICU) frequently experience uncertainty, emotional distress, and significant informational needs. Communication gaps remain common in ICU practice due to time constraints and clinical complexity. Large language models (LLMs) may offer scalable support for patient-family information delivery; however, their performance in responding to real-world ICU family questions has not been systematically evaluated. Methods: This evaluator-blinded, cross-sectional study compared the accuracy of responses generated by five widely used LLMs (Claude Sonnet 4.0, ChatGPT 5.0, Gemini 2.5, Grok-4, and Sonar) to questions commonly asked by ICU family members. A standardized set of 25 questions was generated by prompting each model to list frequently asked ICU family questions. All questions were subsequently posed to all five models in blinded, independent sessions. Two intensive care medicine specialists independently rated response accuracy using a 6-point Likert scale. Inter-rater reliability was assessed using Cohen's kappa. Differences between models were analyzed using the Friedman test with post-hoc Wilcoxon signed-rank tests. Results: A total of 125 responses were evaluated. Inter-rater agreement was moderate (Cohen's kappa = 0.56; overall agreement 73.6%). Accuracy scores differed significantly among models (p < 0.001). Claude Sonnet 4.0 achieved the highest mean accuracy score (5.66 +/- 0.61), followed by ChatGPT 5.0, Gemini 2.5, and Sonar, with no statistically significant differences among these four models. Grok-4 demonstrated significantly lower accuracy compared with all other models (all p < 0.001). Conclusions: Most contemporary LLMs demonstrated high accuracy in answering questions commonly posed by ICU family members, although performance varied across platforms. Selected LLMs may serve as supportive tools to reinforce clinician-family communication; however, careful model selection, clinical oversight, and ethical safeguards are required before implementation in high-stakes intensive care settings.