返回
Beyond rankings: Learning (more) from algorithm validation
DOI:10.1016/j.media.2023.102765.png)
摘要
En 中文
Challenges have become the state-of-the-art approach to benchmark image analysis algorithms in a comparative manner. While the validation on identical data sets was a great step forward, results analysis is often restricted to pure ranking tables, leaving relevant questions unanswered. Specifically, little effort has been put into the systematic investigation on what characterizes images in which state-of-the-art algorithms fail. To address this gap in the literature, we (1) present a statistical framework for learning from challenges and (2) instantiate it for the specific task of instrument instance segmentation in laparoscopic videos. Our framework relies on the semantic meta data annotation of images, which serves as foundation for a General Linear Mixed Models (GLMM) analysis. Based on 51,542 meta data annotations performed on 2,728 images, we applied our approach to the results of the Robust Medical Instrument Segmentation Challenge (ROBUST-MIS) challenge 2019 and revealed underexposure, motion and occlusion of instruments as well as the presence of smoke or other objects in the background as major sources of algorithm failure. Our subsequent method development, tailored to the specific remaining issues, yielded a deep learning model with state-of-the-art overall performance and specific strengths in the processing of images in which previous methods tended to fail. Due to the objectivity and generic applicability of our approach, it could become a valuable tool for validation in the field of medical image analysis and beyond.
Keyword:
Surgical data science
Image characteristics driven algorithm
development
Minimally invasive surgery
Endoscopic vision
Grand challenges
Biomedical image analysis challenges
Generalized linear mixed models
期刊
IF:
11.8
论文数:
3.8K
被引数:
2.4W
机构
引用论文
The influence of various precursors on solar-light-driven g-C3N4 synthesis and its effect on photocatalytic tetracycline hydrochloride (TCH) degradation各种前体对太阳光驱动的g-C3N4合成及其对光催化盐酸四环素 (TCH) 降解的影响
A deep learning framework for quality assessment and restoration in video endoscopy
MEDICAL IMAGE ANALYSIS
IF11.8
BIAS: Transparent reporting of biomedical image analysis challengesBIAS: 生物医学图像分析挑战的透明报告
MEDICAL IMAGE ANALYSIS
IF11.8
Why rankings of biomedical image analysis competitions should be interpreted with care为什么生物医学图像分析比赛的排名应谨慎解释
NATURE COMMUNICATIONS
IF15.7
Using the Unified Protocol for Transdiagnostic Treatment of Emotional Disorders With Youth Exhibiting Anger and Irritability使用统一的协议对表现出愤怒和易怒的年轻人进行情绪障碍的诊断治疗
Deep learning for detection and segmentation of artefact and disease instances in gastrointestinal endoscopy ?深度学习用于检测和分割胃肠内窥镜中的伪道和疾病实例?
MEDICAL IMAGE ANALYSIS
IF11.8

