Language-dependent variation in observed mechanistic performance of web-enabled large language models across myeloid–mucosal immune contexts: a blinded expert evaluation - Takeaways - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Language-Dependent Differences in Mechanistic Performance of Web-Enabled Large Language Models in Myeloid-Mucosal Immune Contexts: An Expert Blinded Assessment

  • By

  • Jinxi Hu

  • Tong Mei

  • Lei Lin

  • Chengyan Xia

  • Tian Ma

  • Penghui Xiang

  • Wenzhuo Wang

  • Siqing Kong

  • Buorui Zhang

  • Sen Cao

  • Yutong Ge

  • Xiaochen Li

  • Lujie Zhang

  • Liangjing Xia

  • Hongliang Duan

  • Chengla Yi

  • September 15, 2026

Share

  • 1

    The study assessed language-dependent differences in mechanistic performance of five web-enabled large language models (LLMs) across four languages.

  • 2

    Twenty senior clinician-experts scored 480 responses using the Mechanistic Fidelity Score, demonstrating moderate-to-good inter-rater agreement.

  • 3

    Reviewer-standardized scores indicated a significant interaction between model and language, but no significant model-by-complexity interaction was found.

  • 4

    Kimi K3 and ChatGPT 5.6 Sol achieved the highest mean Mechanistic Fidelity Scores, significantly outperforming other models.

  • 5

    The results suggest that multilingual biomedical evaluations should directly assess mechanistic explanations in the intended language and context.

Original Source(s)

Related Content