Clinical Scorecard: Evaluating the Performance of AI Agent Systems in Clinical Decision-Making Tasks
At a Glance
Category
Detail
Condition
Clinical decision-making support using agentic AI systems
Key Mechanisms
Autonomous reasoning, planning, and tool invocation by AI agents (e.g., web browsing, code execution, text editing)
Target Population
Healthcare providers and clinical decision support environments
Care Setting
Simulated clinical environments and medical knowledge assessment benchmarks
Key Highlights
Agentic AI systems (OpenManus and Manus) show modest accuracy improvements over baseline large language models in clinical tasks.
Performance varies across benchmarks: up to 60.3% accuracy in diagnostic simulations but low multimodal accuracy (15.5%-29.2%).
Significant computational costs observed, including >10× token usage and >2× latency, with persistent hallucination issues despite safeguards.
Guideline-Based Recommendations
Diagnosis
Use agentic AI systems cautiously as adjuncts given their modest accuracy and hallucination risks.
Validate AI-generated clinical decisions with expert human oversight.
Management
Incorporate AI tools that balance performance gains with computational resource demands.
Prioritize development of more accurate and efficient agentic AI designs for clinical viability.
Monitoring & Follow-up
Continuously monitor AI outputs for hallucinations and errors despite in-agent safeguards.
Track latency and resource usage to ensure integration feasibility in clinical workflows.
Risks
High computational and workflow costs may limit practical deployment.
Persistent hallucinations pose risks for clinical decision accuracy.
Patient & Prescribing Data
Not applicable; study focuses on AI system performance in clinical decision tasks rather than direct patient treatment.
Current agentic AI systems provide limited direct clinical decision support benefits and require further refinement before routine clinical prescribing use.
Clinical Best Practices
Employ agentic AI systems as supplementary tools with human expert validation.
Be aware of the limitations in multimodal data interpretation and accuracy.
Consider computational resource implications when integrating AI agents into clinical workflows.
by Yunsong Liu, Zunamys I. Carrero, Xiaofeng Jiang, Dyke Ferber, Georg Wölflein, Li Zhang, Sanddhya Jayabalan, Tim Lenz, Zhouguang Hui, Jakob Nikolas Kather