AI

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

Researchers have developed a context-fusion framework to improve the accuracy of endoscopic polyp reporting. The framework specializes a pre-trained general-purpose vision-language model (VLM) by introducing specialist knowledge without modifying its weights. This is achieved through implicit instruction context and explicit transduction context, which are used to retrieve related image-report pairs and provide query-specific evidence. Experiments on 2,056 expert-annotated im
Researchers have developed a context-fusion framework to improve the accuracy of endoscopic polyp reporting. The framework specializes a pre-trained general-purpose vision-language model (VLM) by introducing specialist knowledge without modifying its weights. This is achieved through implicit instruction context and explicit transduction context, which are used to retrieve related image-report pairs and provide query-specific evidence. Experiments on 2,056 expert-annotated images showed that the framework outperformed other methods in terms of specialist performance, unified reporting, and adaptation efficiency. --- Why it matters: This matters because it provides a lightweight and effective way to adapt pre-trained VLMs for specialist tasks, which can improve the accuracy of medical image analysis and reporting. This has potential applications in various fields, including healthcare and medical research. Source: https://arxiv.org/abs/2608.15580

This article was originally published at: https://arxiv.org/abs/2608.15580