To systematically evaluate the performance of two AI agent systems in clinical decision-making tasks, highlighting the importance of such evaluations in improving healthcare outcomes.
Key Findings:
OpenManus and Manus achieved modest accuracy gains over baseline LLMs (specifically, the models used for comparison), with scores of 60.3% and 28.0% in AgentClinic MedQA and MIMIC, respectively.
Performance on MedAgentsBench was 30.3%, and only 8.6% on HLE text.
Multimodal accuracy was low, with 15.5% on multimodal HLE and 29.2% on AgentClinic NEJM.
Resource demands increased significantly, with over 10× token usage and more than 2× latency.
89.9% of hallucinations were filtered by in-agent safeguards, but hallucinations remained prevalent.
Interpretation:
Current agentic AI designs provide limited performance benefits while incurring substantial computational and workflow costs, indicating a need for improvement in accuracy, efficiency, and the integration of user feedback.
Limitations:
The study's findings may not generalize to all clinical scenarios due to potential biases in the benchmark families.
The evaluation was limited to specific benchmark families and may not encompass all aspects of clinical decision-making.
Conclusion:
The findings highlight the necessity for more accurate, efficient, and clinically viable AI agent systems in healthcare, suggesting directions for future research.
by Yunsong Liu, Zunamys I. Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, Jakob Nikolas Kather