Oral abstract presentation times

11th Rapid Fire


Abstract ID: 26-241

From Photograph to Surgical Plan: A Comparative Evaluation of Multimodal Artificial Intelligence Systems in the Diagnosis and Surgical Planning of Eyelid Disorders

Author: Ahmet Alp Bilgic
Base Hospital / Institution: Ulucanlar Eye Training and Research Hospital Ankara / Türkiye

Presentation Type: Rapid Fire Presentation
Session: Eyelid Functional
Date: 11th September
Time: 17.56

Purpose

To compare the diagnostic and surgical performance of multimodal artificial intelligence (AI) systems in oculoplastic conditions using photographs alone, and to assess their ability to prioritise the dominant surgical indication in complex cases.


Methods

A total of 246 preoperative frontal photographs of patients with dermatochalasis, ptosis, or eyelid malposition (ectropion or entropion) were analysed; 25 patients underwent combined surgery, defined as simultaneous ptosis correction and upper eyelid blepharoplasty. ChatGPT (GPT-5.3), Gemini 3.0 Flash, and Claude 4.6 Sonnet were evaluated under standardised conditions without clinical context using a fixed prompt. AI-generated diagnoses and surgical recommendations were compared with the surgeon-defined reference standard. Accuracy, Cohen’s kappa, subgroup performance, combined-case prioritisation, and associated finding detection were assessed.


Results

Overall accuracy was similar across models (ChatGPT 74.0%, Gemini 76.4%, Claude 73.2%; all p>0.017). Performance varied by diagnosis. Gemini was most accurate in ptosis (87.0%) and malposition (91.7%), while Claude performed best in dermatochalasis (88.7%) but poorly in ptosis (38.9%), indicating bias. In combined cases (n=25), accuracy declined for ChatGPT (77.8% to 40.0%, p=0.0002) and Claude (79.2% to 20.0%, p<0.001), but remained stable for Gemini (80.0%, p=0.806). The primary failure was prioritisation rather than non-detection: ptosis was recognised as an associated finding in 48% of ChatGPT and 68% of Claude errors, yet not selected as the primary diagnosis. Associated findings were often over-reported.


Conclusion

AI shows moderate performance in oculoplastic decision-making, but accuracy depends on clinical complexity. The main limitation is impaired prioritisation in multi-pathology cases, which may affect surgical planning. These systems show potential as screening and educational tools, but are not ready for autonomous oculoplastic decision-making, and future studies with domain-specific training are warranted.


Additional Authors

First name Last name Base Hospital / Institution
Rukiye Kilic Ucgul Ulucanlar Eye Training and Research Hospital Ankara / Türkiye

↑ Back to top

11th Oral

12th Rapid Fire

12th Oral