Oral abstract presentation times
11th Rapid Fire
Abstract ID: 26-241
From Photograph to Surgical Plan: A Comparative Evaluation of Multimodal Artificial Intelligence Systems in the Diagnosis and Surgical Planning of Eyelid Disorders
Author: Ahmet Alp Bilgic Base Hospital / Institution: Ulucanlar Eye Training and Research Hospital Ankara / Türkiye
Presentation Type: Rapid Fire Presentation Session: Eyelid FunctionalDate: 11th SeptemberTime: 17.56
Purpose
To compare the diagnostic and surgical performance of multimodal artificial intelligence (AI) systems in oculoplastic conditions using photographs alone, and to assess their ability to prioritise the dominant surgical indication in complex cases.
Methods
A total of 246 preoperative frontal photographs of patients with dermatochalasis, ptosis, or eyelid malposition (ectropion or entropion) were analysed; 25 patients underwent combined surgery, defined as simultaneous ptosis correction and upper eyelid blepharoplasty. ChatGPT (GPT-5.3), Gemini 3.0 Flash, and Claude 4.6 Sonnet were evaluated under standardised conditions without clinical context using a fixed prompt. AI-generated diagnoses and surgical recommendations were compared with the surgeon-defined reference standard. Accuracy, Cohen’s kappa, subgroup performance, combined-case prioritisation, and associated finding detection were assessed.
Results
Overall accuracy was similar across models (ChatGPT 74.0%, Gemini 76.4%, Claude 73.2%; all p>0.017). Performance varied by diagnosis. Gemini was most accurate in ptosis (87.0%) and malposition (91.7%), while Claude performed best in dermatochalasis (88.7%) but poorly in ptosis (38.9%), indicating bias. In combined cases (n=25), accuracy declined for ChatGPT (77.8% to 40.0%, p=0.0002) and Claude (79.2% to 20.0%, p<0.001), but remained stable for Gemini (80.0%, p=0.806). The primary failure was prioritisation rather than non-detection: ptosis was recognised as an associated finding in 48% of ChatGPT and 68% of Claude errors, yet not selected as the primary diagnosis. Associated findings were often over-reported.
Conclusion
AI shows moderate performance in oculoplastic decision-making, but accuracy depends on clinical complexity. The main limitation is impaired prioritisation in multi-pathology cases, which may affect surgical planning. These systems show potential as screening and educational tools, but are not ready for autonomous oculoplastic decision-making, and future studies with domain-specific training are warranted.
Additional Authors
| First name | Last name | Base Hospital / Institution |
|---|---|---|
| Rukiye | Kilic Ucgul | Ulucanlar Eye Training and Research Hospital Ankara / Türkiye |
