Evaluating AI Agent Systems in Clinical Decision-Making Tasks
Overview
This study systematically benchmarks two agentic AI systems, OpenManus and Manus, across multiple clinical decision-making tasks. Results show only modest accuracy improvements over baseline large language models, with significant increases in computational resource use and persistent hallucination issues.
Background
Agentic AI systems are designed to autonomously reason, plan, and utilize tools to assist clinical decision-making. Despite their potential, there has been limited systematic evaluation of their real-world performance in healthcare. This study assesses two such systems using diverse benchmark datasets that simulate diagnostic reasoning, medical question answering, and multimodal clinical challenges. Understanding their performance and limitations is critical for developing clinically viable AI tools.
Data Highlights
Benchmark
OpenManus Accuracy
Manus Accuracy
Baseline LLM Accuracy
Resource Usage
AgentClinic MedQA
60.3%
28.0%
Not specified
>10× token usage, >2× latency
MedAgentsBench
30.3%
Not specified
Not specified
HLE Text-only
8.6%
Not specified
Not specified
HLE Multimodal
15.5%
Not specified
Not specified
AgentClinic NEJM Multimodal
29.2%
Not specified
Not specified
Key Findings
Agentic AI systems achieved modest accuracy gains over baseline large language models, with top performance of 60.3% on AgentClinic MedQA.
Multimodal task accuracy remained low, with 15.5% on HLE multimodal and 29.2% on AgentClinic NEJM datasets.
Computational resource demands increased substantially, including over 10-fold token usage and more than double latency compared to baseline models.
Despite in-agent safeguards filtering 89.9% of hallucinations, hallucination prevalence remained a significant issue.
Proprietary Manus system and open-source OpenManus showed variable performance across benchmarks, highlighting challenges in generalizability.
Clinical Implications
Current agentic AI systems provide only limited improvements in clinical decision-making accuracy while imposing significant computational and workflow burdens. Clinicians and healthcare organizations should be cautious in adopting these tools until further advancements improve their reliability and efficiency. Continued development is needed to reduce hallucinations and enhance multimodal reasoning capabilities for practical clinical use.
Conclusion
Agentic AI systems demonstrate potential but currently offer modest clinical performance benefits at high computational cost and with persistent hallucination challenges. Future research must focus on improving accuracy, efficiency, and clinical viability to realize their promise in healthcare.
by Yunsong Liu, Zunamys I. Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, Jakob Nikolas Kather