Benchmarking large language model-based agent systems for clinical decision tasks - Summary - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Evaluating the Performance of AI Agent Systems in Clinical Decision-Making Tasks

  • By

  • Yunsong Liu

  • Zunamys I. Carrero

  • Xiaofeng Jiang

  • Dyke Ferber

  • Georg Wölflein

  • Li Zhang

  • Sanddhya Jayabalan

  • Tim Lenz

  • Zhouguang Hui

  • Jakob Nikolas Kather

  • February 18, 2026

Share

Objective:

To systematically evaluate the performance of two AI agent systems in clinical decision-making tasks, highlighting the importance of such evaluations in improving healthcare outcomes.

Key Findings:
  • OpenManus and Manus achieved modest accuracy gains over baseline LLMs (specifically, the models used for comparison), with scores of 60.3% and 28.0% in AgentClinic MedQA and MIMIC, respectively.
  • Performance on MedAgentsBench was 30.3%, and only 8.6% on HLE text.
  • Multimodal accuracy was low, with 15.5% on multimodal HLE and 29.2% on AgentClinic NEJM.
  • Resource demands increased significantly, with over 10× token usage and more than 2× latency.
  • 89.9% of hallucinations were filtered by in-agent safeguards, but hallucinations remained prevalent.
Interpretation:

Current agentic AI designs provide limited performance benefits while incurring substantial computational and workflow costs, indicating a need for improvement in accuracy, efficiency, and the integration of user feedback.

Limitations:
  • The study's findings may not generalize to all clinical scenarios due to potential biases in the benchmark families.
  • The evaluation was limited to specific benchmark families and may not encompass all aspects of clinical decision-making.
Conclusion:

The findings highlight the necessity for more accurate, efficient, and clinically viable AI agent systems in healthcare, suggesting directions for future research.

Original Source(s)

Related Content