Abstract ID: 26-243

Evaluating artificial intelligence chatbot quality for thyroid eye disease patient information using validated assessment tools

Author: Hamad Hejazi
Base Hospital / Institution: Newcastle upon Tyne Hospitals NHS Foundation Trust

Presentation Type: ePoster Presentation

Purpose

Patients with Thyroid Eye Disease use AI chatbots to obtain information about their condition. This study aims to evaluate the accuracy, readability and quality of responses generated by 3 leading AI LLMs as they respond to frequently asked questions about TED. This study uses validated assessment frameworks.


Methods

Twelve questions representative of common TED patient queries were identified from published literature and clinical experience, spanning pathophysiology, symptoms, treatment options, and surgical management. Questions were independently submitted to GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet. Responses were evaluated using the DISCERN instrument (scored 1-5 across 16 items) and Flesch-Kincaid Grade Level for readability. Two independent raters scored all responses; inter-rater reliability was assessed using intraclass correlation coefficient (ICC).


Results

Results pending at time of submission. Responses will be assessed across all three models for overall DISCERN score, subscale performance, and readability. Statistical comparison will employ Kruskal-Wallis testing with post-hoc analysis. Preliminary analysis is underway.


Conclusion

To our knowledge, this is the first study to apply the validated DISCERN instrument to AI-generated TED patient information. Findings will identify variation in quality across platforms, inform clinical guidance on AI tool use, and highlight areas where current LLMs may mislead patients seeking information about this complex condition.


Additional Authors

First name Last name Base Hospital / Institution
Abdullah Sheekhuna Gateshead NHS Trust

↑ Back to top