Abstract ID: 26-243
Evaluating artificial intelligence chatbot quality for thyroid eye disease patient information using validated assessment tools
Author: Hamad Hejazi Base Hospital / Institution: Newcastle upon Tyne Hospitals NHS Foundation Trust
Presentation Type: ePoster Presentation
Purpose
Patients with Thyroid Eye Disease use AI chatbots to obtain information about their condition. This study aims to evaluate the accuracy, readability and quality of responses generated by 3 leading AI LLMs as they respond to frequently asked questions about TED. This study uses validated assessment frameworks.
Methods
Twelve questions representative of common TED patient queries were identified from published literature and clinical experience, spanning pathophysiology, symptoms, treatment options, and surgical management. Questions were independently submitted to GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Responses were evaluated using the DISCERN instrument (scored 1-5 across 16 items) and Flesch-Kincaid Grade Level for readability. Two independent raters scored all responses; inter-rater reliability was assessed using intraclass correlation coefficient (ICC).
Results
Results pending at time of submission. Responses will be assessed across all three models for overall DISCERN score, subscale performance, and readability. Statistical comparison will employ Kruskal-Wallis testing with post-hoc analysis. Preliminary analysis is underway.
Conclusion
To our knowledge, this is the first study to apply the validated DISCERN instrument to AI-generated TED patient information. Findings will identify variation in quality across platforms, inform clinical guidance on AI tool use, and highlight areas where current LLMs may mislead patients seeking information about this complex condition.
Additional Authors
| First name | Last name | Base Hospital / Institution |
|---|---|---|
| Abdullah | Sheekhuna | Gateshead NHS Trust |
