Halil İbrahim UZUNLU, Gokce YILDIRAN, Hande AKDENIZ, Zekeriya TOSUN
Turkish Journal of Plastic Surgery - 2026;34(2):51-56
Background: In daily practice, plastic surgeons frequently distinguish between benign and malignant skin lesions to plan appropriate excisions. Artificial intelligence (AI) algorithms offer potential decision support in this domain. This study evaluated the diagnostic accuracy of ModelDerm, a deep-learning algorithm, using real-world clinical data from a tertiary plastic surgery center. Materials and Methods: This retrospective study included 354 patients (total 378 images) presenting with cutaneous lesions. Clinical photographs were analyzed by ModelDerm, which provides five ranked diagnoses and a malignancy risk score. The AI outputs were compared with the gold-standard histopathological diagnoses. Performance was assessed using match rates, sensitivity, specificity, and the area under the receiver operating characteristic curve (AUC). Results: Of the 104 malignant and 250 benign lesions analyzed, the correct diagnosis appeared within the AI's top-5 candidates in 55.4% of cases. The correct diagnosis was ranked first in 29.7% of cases. For differentiating benign from malignant lesions, the algorithm achieved an AUC of 0.89 (95% confidence interval: 0.85-0.93). The sensitivity for malignancy detection (rank-one output "skin cancer") was 82.7% (86/104 cases), and specificity was 93.6% (234/250 cases). Performance was particularly robust in Fitzpatrick skin types III and IV (AUC 0.92 and 0.98, respectively). Conclusions: ModelDerm demonstrated strong discriminative ability between benign and malignant lesions. While the system shows promise as a clinical decision-support tool to reduce diagnostic uncertainty, particularly for triaging suspicious lesions, human oversight remains essential due to limitations in specific subtype diagnosis.