To enhance safety in surgical Visual Question Answering (VQA) by developing a failure detection method that incorporates question alignment into uncertainty estimation for preclinical evaluation.
Approach:
QA-SNNE Development: Introduced question-aligned semantic nearest neighbor entropy (QA-SNNE) as a black-box failure-detection score that combines answer consistency with question alignment.
Out-of-template Evaluation: Constructed an out-of-template version of EndoVis18-VQA by rephrasing question templates while preserving images, answers, and splits to test model stability under varied wording.
Evaluation Metrics: Evaluated QA-SNNE against existing methods (DSE, SNNE, VL-U) using AUROC, calibrated accuracy, sensitivity, and specificity across zero-shot and PEFT surgical VQA models.
Key Findings:
QA-SNNE effectively captures both answer consistency and question validity, addressing limitations of existing methods.
The out-of-template evaluation revealed that models may not perform reliably under varied question phrasing.
QA-SNNE demonstrated improved failure detection compared to traditional methods.
Interpretation:
The study highlights the integration of question alignment into uncertainty estimation for surgical VQA.
Limitations:
The study does not claim clinical validation or routine deployment of the proposed methods.
Evaluation is based on a controlled dataset, which may not fully represent real-world variability in clinical language.
Conclusion:
QA-SNNE improves the reliability of surgical VQA systems by addressing both answer consistency and question relevance.