Abstract ID: 26-440
Can Large Language Models Recognize Benign Eyelid Lesions? A Preliminary Pathology-Based Study
Author: Dilara berrin Guzel Base Hospital / Institution: Kanuni Sultan Süleyman Training and Research Hospital
Presentation Type: Rapid Fire Presentation Session: Ptosis & MiscellaneousDate: 12th SeptemberTime: 09.15AM
Purpose
To compare the performance of an ophthalmologist and two multimodal large language models (LLMs), ChatGPT and Gemini, in benign eyelid lesions, focusing on benign-malignant discrimination and histopathological diagnosis prediction.
Methods
This retrospective preliminary study included 110 patients with histopathologically confirmed benign eyelid lesions who underwent surgical excision. Preoperative clinical photographs were evaluated by an ophthalmologist. The same images were subsequently assessed by ChatGPT and Gemini using a standardized prompt. Evaluators classified lesions as benign or malignant/suspicious and predicted the most likely histopathological diagnosis. For LLMs, the top five differential diagnoses were additionally recorded when the exact diagnosis was incorrect. Paired comparisons were performed using McNemar tests.
Results
The mean age was 49.7 ± 16.8 years (range, 2-85 years), and 67 patients (60.9%) were female. False-positive malignancy suspicion rates were identical for the ophthalmologist, ChatGPT, and Gemini (6/110, 5.5% each), with no significant differences between evaluators (all p = 1.000). Histopathological diagnosis accuracy was 61.8% for the ophthalmologist, 34.5% for ChatGPT, and 36.4% for Gemini. The ophthalmologist demonstrated significantly higher accuracy than both ChatGPT and Gemini (both p < 0.001), whereas no significant difference was observed between the two LLMs (p = 0.655). The correct diagnosis was included among the top five differential diagnoses in 56.4% of cases for both LLMs.
Conclusion
Although LLMs showed low false-positive malignancy suspicion rates in benign eyelid lesions, their histopathological diagnostic accuracy remained substantially lower than that of the ophthalmologist. Nevertheless, inclusion of the correct diagnosis among the top five differential diagnoses suggests a potential supportive role for LLMs in eyelid lesion assessment.
Additional Authors
| First name | Last name | Base Hospital / Institution |
|---|---|---|
| Ece yalcindag | Aydin | Kanuni sultan suleyman training and research hospital |
| Fatma esin | Ozdemir | Kanuni sultan Süleyman training and research hospital |
