Toward Structural Multimodal Representations: Specialization, Selection, and Sparsification via Mixture-of-Experts
arXiv:2605.03348v2 Announce Type: replace
Abstract: We propose S3 (Specialization, Selection, Sparsification), a framework that rethinks multimodal learning through a structural perspective. Instead of encoding all signals into a fixed embedding, S3 d…