Source: arXiv · cs.CVView original ↗
Copyright remains with the original source. This site only collects, translates, or reformats the material.
What happened
arXiv:2608.07543v1 Announce Type: new Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showin
Analysis and impact
结肠息肉的准确光学诊断,是决定内镜下切除方式与术后随访间隔的关键环节。传统上,这一任务高度依赖内镜医师的实时判断与经验积累。随着多模态大语言模型的快速发展,研究者开始探索这些能够同时理解图像与文本的模型,能否在这一精细的医学影像分析场景中提供可靠支持。
该研究聚焦于评估多模态大语言模型在结肠息肉光学诊断中的性能。研究团队通过设计特定的测试方案,将模型对内镜图像的解读结果,与病理金标准或专家诊断进行对比。摘要指出,准确的诊断对于指导切除策略和后续监测至关重要,这构成了评估模型临床价值的核心前提。原文未提供具体的模型名称、测试数据集规模及详细的性能指标数据。
从技术背景看,多模态大语言模型在医学影像领域的应用正成为研究热点。与仅输出单一建议的传统辅助诊断系统不同,这类模型具备生成描述性报告、回答针对性问题的能力,其交互方式更贴近临床咨询场景。不过,将其部署于息肉光学诊断等实时、高风险任务,仍需克服模型幻觉、细粒度特征识别准确性以及对罕见亚型泛化能力等挑战。该研究为理解当前大模型在这一特定内镜子任务中的能力边界提供了实证参考。
References
Original source text
arXiv:2608.07543v1 Announce Type: new Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.