Abdullah Enes ATAŞ, Halil İbrahim ŞARA, Fatih Cemal TEKIN
Eurasian Journal of Emergency Medicine - 2026;25(1):411-418
Aim: To evaluate the accuracy of four large language models in assessing the appropriateness of imaging tests for patients with suspected acute abdomen, using the American College of Radiology (ACR) Appropriateness Criteria as the gold standard. Materials and Methods: This cross-sectional study evaluated ChatGPT 5.5, Gemini 3.1 Pro, Claude Opus 4.8, and DeepSeek V4 Pro across 30 clinical scenarios of acute abdominal pain. Models were tested using zero-shot and one-shot prompting strategies, yielding a total of 240 responses. Responses were compared to the ACR categories of "Usually Appropriate, "May Be Appropriate," and "Usually Not Appropriate." Performance was measured using exact match accuracy, linearly weighted Cohen's kappa, and major discordance rates. Results: The overall accuracy of the pooled models was 86.3%, with a weighted kappa of 0.78. Claude Opus 4.8 achieved the highest accuracy at 95.0% and a kappa of 0.91. Clinically risky major discordance rates remained low across all models, ranging from 3.3% to 8.3%. Utilizing a one-shot prompting strategy increased overall accuracy from 81.7% to 90.8% compared to zero-shot prompting. Conclusion: Large language models successfully align imaging recommendations for acute abdominal conditions with established ACR guidelines. Providing a guideline-formatted example prior to the task improves model performance and highlights their potential as reliable radiological decision-support tools. By supporting faster, guideline-concordant imaging selection, such tools may help reduce unnecessary ionizing radiation exposure and alleviate decision-making pressure in high-volume emergency settings.