Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
arXiv:2603.25155v1 Announce Type: new
Abstract: Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fi…