To examine the dimensions of operational and decisional trust in AI systems for clinical decision support and to evaluate the performance of an on-premise AI agent.
Approach:
Operational Trust: Implemented a clinically governable AI system and measured end-to-end performance with competitive open-weight LLMs.
Decisional Trust: Developed a multi-perspective confidence framework to assess reliability through internal likelihood, linguistic expression of uncertainty, and behavioral stability.
Key Findings:
The on-premise dual-agent framework achieved competitive diagnostic performance with Qwen-3.5 recording 90.0% accuracy on the MIRA-v2 benchmark.
The framework demonstrated the ability to differentiate between cases suitable for autonomous handling and those requiring clinician review.
Confidence estimation and calibration are crucial for ensuring safety and reliability in AI outputs.
Interpretation:
The study examines the importance of operational and decisional trust in deploying AI systems in clinical settings, highlighting the need for reliable signals during decision-making.
Limitations:
Current methods for confidence estimation are minimally integrated into agent evaluation.
The evaluation primarily focused on specific benchmarks and may not generalize to all clinical scenarios.
Conclusion:
A framework for assessing confidence signals in clinical AI systems is discussed to support selective autonomy while ensuring governance.
by Li Zhang, Georg Wölflein, Dyke Ferber, Junhao Liang, Zunamys I. Carrero, Xuewei Wu, Julien Vibert, Jan Clusmann, Lino Möhrmann, Elena E. Möhrmann, Catharina Wichmann, Fabian Wolf, Tim Lenz, Jakob Nikolas Kather