Patient-facing diabetic foot information from large language models: a domain- and source-balanced prompt framework for public-interface benchmarking - Summary - MDSpire
Advertisement
Evaluating Patient-Centric Diabetic Foot Information Generated by Large Language Models: A Framework for Balanced Prompt Design and Public Interface Assessment
To develop and apply a domain- and source-balanced prompt framework for benchmarking patient-facing diabetic foot information generated by publicly accessible LLMs.
Approach:
Benchmark Prompt Set: A 24-item benchmark prompt set was generated using a domain- and source-balanced framework incorporating public-query sources and guideline-derived decision-critical content.
Model Evaluation: Responses from five LLMs (GPT-5.5 Thinking, DeepSeek-V4, Gemini 3.1 Pro, Grok 4.3, and Qwen3.6-Max-Preview) were assessed for quality, transparency, readability, and clinical-risk signals using established evaluation criteria.
Key Findings:
Grok 4.3 recorded the highest observed mean scores for DISCERN, EQIP, GQS, and JAMA-based transparency-related metrics.
DeepSeek-V4 showed the lowest mean scores for several readability-grade metrics.
No response was rated as PCF 1 or PCF 2, indicating no overt short-term harm signals.
No response met all predefined readability targets.
Interpretation:
The proposed prompt framework provides a structured basis for public-interface LLM benchmarking in diabetic foot education.
Limitations:
Formal claim-level factual-accuracy review and guideline-concordance adjudication were not performed.
Responses were evaluated under default public-interface conditions, which may not reflect real-world usage.
Conclusion:
Default responses showed metric-specific variation and limited readability.
A 20-week INSIGHT extension suggests that switching from reference aflibercept to MYL-1701P maintained comparable safety, visual, and anatomic outcomes in patients with diabetic macular edema.