Benchmarking large language model-based agent systems for clinical decision tasks - Scorecard - MDSpire
Coming Soon: Introducing MDSpire News. Learn more
Conexiant’s news site is now MDSpire News. Learn more

Evaluating the Performance of AI Agent Systems in Clinical Decision-Making Tasks

  • By

  • Yunsong Liu

  • Zunamys I. Carrero

  • Xiaofeng Jiang

  • Dyke Ferber

  • Georg Wölflein

  • Li Zhang

  • Sanddhya Jayabalan

  • Tim Lenz

  • Zhouguang Hui

  • Jakob Nikolas Kather

  • February 18, 2026

Share

Clinical Scorecard: Evaluating the Performance of AI Agent Systems in Clinical Decision-Making Tasks

At a Glance

CategoryDetail
ConditionClinical decision-making support using agentic AI systems
Key MechanismsAutonomous reasoning, planning, and tool invocation by AI agents (e.g., web browsing, code execution, text editing)
Target PopulationHealthcare providers and clinical decision support environments
Care SettingSimulated clinical environments and medical knowledge assessment benchmarks

Key Highlights

  • Agentic AI systems (OpenManus and Manus) show modest accuracy improvements over baseline large language models in clinical tasks.
  • Performance varies across benchmarks: up to 60.3% accuracy in diagnostic simulations but low multimodal accuracy (15.5%-29.2%).
  • Significant computational costs observed, including >10× token usage and >2× latency, with persistent hallucination issues despite safeguards.

Guideline-Based Recommendations

Diagnosis

  • Use agentic AI systems cautiously as adjuncts given their modest accuracy and hallucination risks.
  • Validate AI-generated clinical decisions with expert human oversight.

Management

  • Incorporate AI tools that balance performance gains with computational resource demands.
  • Prioritize development of more accurate and efficient agentic AI designs for clinical viability.

Monitoring & Follow-up

  • Continuously monitor AI outputs for hallucinations and errors despite in-agent safeguards.
  • Track latency and resource usage to ensure integration feasibility in clinical workflows.

Risks

  • High computational and workflow costs may limit practical deployment.
  • Persistent hallucinations pose risks for clinical decision accuracy.

Patient & Prescribing Data

Not applicable; study focuses on AI system performance in clinical decision tasks rather than direct patient treatment.

Current agentic AI systems provide limited direct clinical decision support benefits and require further refinement before routine clinical prescribing use.

Clinical Best Practices

  • Employ agentic AI systems as supplementary tools with human expert validation.
  • Be aware of the limitations in multimodal data interpretation and accuracy.
  • Consider computational resource implications when integrating AI agents into clinical workflows.

References

Original Source(s)

Related Content