Benchmarking large language model-based agent systems for clinical decision tasks - Report - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Evaluating the Performance of AI Agent Systems in Clinical Decision-Making Tasks

  • By

  • Yunsong Liu

  • Zunamys I. Carrero

  • Xiaofeng Jiang

  • Dyke Ferber

  • Georg Wölflein

  • Li Zhang

  • Sanddhya Jayabalan

  • Tim Lenz

  • Zhouguang Hui

  • Jakob Nikolas Kather

  • February 18, 2026

Share

Evaluating AI Agent Systems in Clinical Decision-Making Tasks

Overview

This study systematically benchmarks two agentic AI systems, OpenManus and Manus, across multiple clinical decision-making tasks. Results show only modest accuracy improvements over baseline large language models, with significant increases in computational resource use and persistent hallucination issues.

Background

Agentic AI systems are designed to autonomously reason, plan, and utilize tools to assist clinical decision-making. Despite their potential, there has been limited systematic evaluation of their real-world performance in healthcare. This study assesses two such systems using diverse benchmark datasets that simulate diagnostic reasoning, medical question answering, and multimodal clinical challenges. Understanding their performance and limitations is critical for developing clinically viable AI tools.

Data Highlights

BenchmarkOpenManus AccuracyManus AccuracyBaseline LLM AccuracyResource Usage
AgentClinic MedQA60.3%28.0%Not specified>10× token usage, >2× latency
MedAgentsBench30.3%Not specifiedNot specified
HLE Text-only8.6%Not specifiedNot specified
HLE Multimodal15.5%Not specifiedNot specified
AgentClinic NEJM Multimodal29.2%Not specifiedNot specified

Key Findings

  • Agentic AI systems achieved modest accuracy gains over baseline large language models, with top performance of 60.3% on AgentClinic MedQA.
  • Multimodal task accuracy remained low, with 15.5% on HLE multimodal and 29.2% on AgentClinic NEJM datasets.
  • Computational resource demands increased substantially, including over 10-fold token usage and more than double latency compared to baseline models.
  • Despite in-agent safeguards filtering 89.9% of hallucinations, hallucination prevalence remained a significant issue.
  • Proprietary Manus system and open-source OpenManus showed variable performance across benchmarks, highlighting challenges in generalizability.

Clinical Implications

Current agentic AI systems provide only limited improvements in clinical decision-making accuracy while imposing significant computational and workflow burdens. Clinicians and healthcare organizations should be cautious in adopting these tools until further advancements improve their reliability and efficiency. Continued development is needed to reduce hallucinations and enhance multimodal reasoning capabilities for practical clinical use.

Conclusion

Agentic AI systems demonstrate potential but currently offer modest clinical performance benefits at high computational cost and with persistent hallucination challenges. Future research must focus on improving accuracy, efficiency, and clinical viability to realize their promise in healthcare.

References

  1. Schmidgall et al. 2024 -- AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
  2. Jiang et al. 2025 -- MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLMAgents
  3. Shortliffe & Sepúlveda 2018 -- Clinical decision support in the era of artificial intelligence
  4. Elhaddad & Hamam 2024 -- AI-driven clinical decision support systems: an ongoing pursuit of potential

Original Source(s)

Related Content