Background: Patients increasingly consult artificial intelligence (AI) chatbots for health information, yet the reliability and accessibility of AI-generated content for complex conditions like dysphagia remain unvalidated. We conducted the first head-to-head comparison of leading large language models for dysphagia patient education, evaluated by an international expert panel. Methods: Forty-six validated questions across four clinical domains were submitted to ChatGPT-4.0 (OpenAI) and Claude 3.7 (Anthropic) in March 2025. Ten blinded experts from six countries rated responses for scientific accuracy (5-point Likert), clarity (5-point Likert), and misinformation (binary). Readability was assessed using Flesch Reading Ease, Flesch–Kincaid Grade Level, and SMOG Index. Between-model comparisons used Wilcoxon signed-rank tests with Cohen's d effect sizes. Key Results: No significant differences emerged for scientific accuracy (ChatGPT: 3.87 ± 0.36 vs. Claude: 3.93 ± 0.35; p = 0.26; d = 0.16), clarity (4.12 ± 0.34 vs. 4.15 ± 0.27; p = 0.67; d = 0.11), or mean misinformation rates (both 2.15; p = 0.96). Strong inter-model correlation existed for accuracy (rs = 0.678; p < 0.001). Critically, both models produced content far exceeding recommended readability levels: SMOG indices of 14.95 ± 2.40 years (ChatGPT) and 17.37 ± 2.67 years (Claude) required extensive education (p < 0.001; d = 0.95) versus the recommended 6–7 years. Categorical analysis showed Claude generated three times more misinformation-free responses (19.6% vs. 6.5%; p = 0.077). Conclusions and Inferences: Leading AI chatbots demonstrate equivalent, acceptable accuracy for dysphagia information but produce content inaccessible to most patients due to excessive complexity. The strong inter-model correlation suggests shared limitations in medical training data. Before clinical implementation, AI-generated patient education requires mandatory readability optimization to address the substantial health literacy gap identified in this study.
Artificial Intelligence Chatbots for Dysphagia Patient Education: A Multi‐Center International Expert Evaluation
Marabotto, Elisa;Savarino, Vincenzo;
2026-01-01
Abstract
Background: Patients increasingly consult artificial intelligence (AI) chatbots for health information, yet the reliability and accessibility of AI-generated content for complex conditions like dysphagia remain unvalidated. We conducted the first head-to-head comparison of leading large language models for dysphagia patient education, evaluated by an international expert panel. Methods: Forty-six validated questions across four clinical domains were submitted to ChatGPT-4.0 (OpenAI) and Claude 3.7 (Anthropic) in March 2025. Ten blinded experts from six countries rated responses for scientific accuracy (5-point Likert), clarity (5-point Likert), and misinformation (binary). Readability was assessed using Flesch Reading Ease, Flesch–Kincaid Grade Level, and SMOG Index. Between-model comparisons used Wilcoxon signed-rank tests with Cohen's d effect sizes. Key Results: No significant differences emerged for scientific accuracy (ChatGPT: 3.87 ± 0.36 vs. Claude: 3.93 ± 0.35; p = 0.26; d = 0.16), clarity (4.12 ± 0.34 vs. 4.15 ± 0.27; p = 0.67; d = 0.11), or mean misinformation rates (both 2.15; p = 0.96). Strong inter-model correlation existed for accuracy (rs = 0.678; p < 0.001). Critically, both models produced content far exceeding recommended readability levels: SMOG indices of 14.95 ± 2.40 years (ChatGPT) and 17.37 ± 2.67 years (Claude) required extensive education (p < 0.001; d = 0.95) versus the recommended 6–7 years. Categorical analysis showed Claude generated three times more misinformation-free responses (19.6% vs. 6.5%; p = 0.077). Conclusions and Inferences: Leading AI chatbots demonstrate equivalent, acceptable accuracy for dysphagia information but produce content inaccessible to most patients due to excessive complexity. The strong inter-model correlation suggests shared limitations in medical training data. Before clinical implementation, AI-generated patient education requires mandatory readability optimization to address the substantial health literacy gap identified in this study.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



