Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs
Published in Neural Information Processing Systems (NeurIPS) (Spotlight · Top 1.3%), 2026
Spotlight · Top 1.3%
Directional Cross-Modal Alignment Transfer (DCAT) transfers textual alignment from a strong, well-aligned source modality to a weaker target modality, improving audio and video MLLMs without further fine-tuning. The method uses a closed-form weight-space update computed from a small calibration set. Theoretical and empirical analyses connect cross-modal alignment to downstream performance, providing a principled basis for efficient transfer across modalities.
Accepted as a Spotlight (Top 1.3%) at NeurIPS 2026. The arXiv preprint is not yet available.
Authors: Hoigi Seo*, Byung Hyun Lee*, Minjun Kim*, Dohyun Mah, Jongho Lee and Se Young Chun. (* co-first author)
