Yasin ETLİ, Erhan KARTAL, Mahmut AŞIRDIZER
Van Medical Journal - 2026;33(3):228-237
Introduction: Disability impairment rating is a critical medicolegal function requiring precise integration of clinical findings with regulatory tables. Large language models (LLMs) have shown promising capabilities in medicine; however, their ability to perform quantitative disability assessments has not been evaluated. This study aimed to assess the agreement between three commercial LLMs and expert-determined disability ratings. Materials and Methods: One hundred standardized fictional clinical cases were constructed representing the spectrum of medicolegal disability assessments. Each case was evaluated by a forensic medicine specialist using the Turkish Regulation on Disability Assessment for Adults (2019) as the reference standard. Three LLMs-GPT-5, Gemini 3 Pro, and Claude Opus 4.6-were tested using a zero-shot prompt. A three-model ensemble was also examined. Agreement was assessed using intraclass correlation coefficients (ICC), Bland-Altman analysis, and mean absolute error (MAE). Results: All three LLMs demonstrated good agreement with the reference standard, with ICC values of 0.875 (Gemini 3 Pro), 0.848 (Claude Opus 4.6), and 0.814 (GPT-5). MAE ranged from 5.1 to 6.0 percentage points. No significant difference was found among models (Friedman p=0.790). The three-model ensemble achieved the highest ICC (0.890) and lowest MAE (4.7 points). Accuracy declined substantially for rates >=21% (MAE 9.1-11.5). Agreement for temporary work disability duration was moderate (ICC 0.701-0.750), while caregiver need showed poor-to-moderate agreement (ICC 0.497-0.585). Conclusion: Commercial LLMs demonstrate good overall agreement with expert disability ratings but show declining accuracy for complex multi-system cases. A multi-model ensemble approach improves reliability. LLMs may serve as supportive tools in medicolegal disability assessment, although expert oversight remains essential.