Publicly Accessible Large Language Model Responses to Frequently Asked Questions About Spondylodiscitis: Preliminary Expert Evaluation - Summary - MDSpire

Evaluation of Responses from Publicly Available Large Language Models to Common Inquiries Regarding Spondylodiscitis: Initial Expert Assessment

  • By

  • Melanie Ardelt

  • David Schiffelholz

  • Siegmund Lang

  • Josina Straub

  • Sonja Häckel

  • Nicolas von der Hoeh

  • Marc Dreimann

  • Jonathan Neuhoff

  • Sebastian Siller

  • Denis Bratelj

  • Volker Alt

  • Dietmar Dammerer

  • Jonas Krueckel

  • July 16, 2026

Share

Objective:

To evaluate spine surgeons’ ratings of responses generated by large language models (LLMs) to frequently asked questions about spondylodiscitis.

Approach:
  • Identification of Relevant FAQs: A chronological, multisource workflow was applied to identify patient-oriented FAQs about spondylodiscitis through Google and PubMed searches.
Key Findings:
  • Spondylodiscitis is a rare but increasingly common condition with high morbidity and an in-hospital mortality rate of 17.2% within the first year after diagnosis.
  • Patients often seek additional information online, where LLMs like ChatGPT and Google Gemini can provide accessible health information.
  • The performance of LLMs in generating responses specific to spondylodiscitis has not been systematically evaluated prior to this study.
Interpretation:

The study highlights the need for evaluating LLMs in the context of complex medical conditions like spondylodiscitis to ensure the accuracy and reliability of the information provided.

Limitations:
  • The study is preliminary and focuses on a limited set of FAQs.
  • Responses were rated by spine surgeons, which may not fully represent patient perspectives.
Conclusion:

This evaluation serves as an initial step in assessing the utility of LLMs for providing information on spondylodiscitis.

Sources:

Original Source(s)

Related Content