Abstract ID: 26-440

Can Large Language Models Recognize Benign Eyelid Lesions? A Preliminary Pathology-Based Study

Author: Dilara berrin Guzel
Base Hospital / Institution: Kanuni Sultan Süleyman Training and Research Hospital

Presentation Type: Rapid Fire Presentation
Session: Ptosis & Miscellaneous
Date: 12th September
Time: 09.15AM

Purpose

To compare the performance of an ophthalmologist and two multimodal large language models (LLMs), ChatGPT and Gemini, in benign eyelid lesions, focusing on benign-malignant discrimination and histopathological diagnosis prediction.


Methods

This retrospective preliminary study included 110 patients with histopathologically confirmed benign eyelid lesions who underwent surgical excision. Preoperative clinical photographs were evaluated by an ophthalmologist. The same images were subsequently assessed by ChatGPT and Gemini using a standardized prompt. Evaluators classified lesions as benign or malignant/suspicious and predicted the most likely histopathological diagnosis. For LLMs, the top five differential diagnoses were additionally recorded when the exact diagnosis was incorrect. Paired comparisons were performed using McNemar tests.


Results

The mean age was 49.7 ± 16.8 years (range, 2-85 years), and 67 patients (60.9%) were female. False-positive malignancy suspicion rates were identical for the ophthalmologist, ChatGPT, and Gemini (6/110, 5.5% each), with no significant differences between evaluators (all p = 1.000). Histopathological diagnosis accuracy was 61.8% for the ophthalmologist, 34.5% for ChatGPT, and 36.4% for Gemini. The ophthalmologist demonstrated significantly higher accuracy than both ChatGPT and Gemini (both p < 0.001), whereas no significant difference was observed between the two LLMs (p = 0.655). The correct diagnosis was included among the top five differential diagnoses in 56.4% of cases for both LLMs.


Conclusion

Although LLMs showed low false-positive malignancy suspicion rates in benign eyelid lesions, their histopathological diagnostic accuracy remained substantially lower than that of the ophthalmologist. Nevertheless, inclusion of the correct diagnosis among the top five differential diagnoses suggests a potential supportive role for LLMs in eyelid lesion assessment.


Additional Authors

First name Last name Base Hospital / Institution
Ece yalcindag Aydin Kanuni sultan suleyman training and research hospital
Fatma esin Ozdemir Kanuni sultan Süleyman training and research hospital

↑ Back to top