To evaluate the performance of an ImageNet-pretrained vision transformer (ViT) in predicting endoscopic and histologic activity in ulcerative colitis (UC) using white-light colonoscopy videos.
Approach:
Methodological Concerns: The study raises concerns about the alignment of the ViT architecture with data representation and clinical objectives, suggesting that the processing may attenuate critical visual signals relevant to UC.
Label Generation Issues: The method of generating labels for histologic healing is critiqued for being weak and spatially imprecise, with suggestions for more robust labeling approaches.
Evaluation Beyond Accuracy: The article emphasizes the need for evaluation metrics beyond accuracy, including model calibration and decision-curve analysis, to assess clinical risks associated with predictions.
Key Findings:
ViTs may not optimally capture clinically decisive endoscopic features due to processing methods.
Labeling for histologic healing is weak and may not accurately reflect the underlying mucosal conditions.
Clinical translation requires comprehensive evaluation metrics that account for clinical risks.
Interpretation:
The study highlights the potential of AI in UC assessment but calls for improvements in model robustness, interpretability, and clinical integration.
Limitations:
The processing of video data may lead to loss of important visual signals.
Label generation methods may introduce noise and spatial discordance.
Current evaluation metrics may not fully capture clinical implications of model predictions.
Conclusion:
Future research should focus on developing AI systems that are robust, interpretable, and integrated into clinical workflows.