Patient-facing diabetic foot information from large language models: a domain- and source-balanced prompt framework for public-interface benchmarking - Report - MDSpire
Advertisement
Evaluating Patient-Centric Diabetic Foot Information Generated by Large Language Models: A Framework for Balanced Prompt Design and Public Interface Assessment
Clinical Report: Evaluating Patient-Centric Diabetic Foot Information Generated by LLMs
Overview
This study developed a framework for benchmarking diabetic foot information generated by large language models (LLMs).
Background
Diabetes-related foot disease is a leading cause of preventable complications, including hospitalization and limb loss. Timely recognition of risk factors and effective patient education are crucial for improving outcomes. The use of LLMs to generate patient-facing information presents an opportunity to enhance understanding.
Data Highlights
No formal numerical data or trial results were provided in the source material.
Key Findings
The study generated a 24-item benchmark prompt set for evaluating LLM outputs.
Grok 4.3 achieved the highest scores in quality metrics such as DISCERN, EQIP, and GQS.
DeepSeek-V4 recorded the lowest readability scores and the highest mean FRES.
No responses met all predefined readability targets.
Visible transparency-related scores were low across all models.
No response was flagged as posing overt short-term harm (PCF 1 or PCF 2).
Clinical Implications
The variability in response quality highlights the need for clinician oversight in patient education derived from these models.
Conclusion
The study reveals significant limitations in readability and transparency in LLM outputs for diabetic foot education.