Strong Helps Weak: Directional Cross-Modal Alignment Transfer in Multi-modal LLMs

Published in Neural Information Processing Systems (NeurIPS) (Spotlight · Top 1.3%), 2026

Spotlight · Top 1.3%

Directional Cross-Modal Alignment Transfer (DCAT) transfers textual alignment from a strong, well-aligned source modality to a weaker target modality, improving audio and video MLLMs without further fine-tuning. The method uses a closed-form weight-space update computed from a small calibration set. Theoretical and empirical analyses connect cross-modal alignment to downstream performance, providing a principled basis for efficient transfer across modalities.

Accepted as a Spotlight (Top 1.3%) at NeurIPS 2026. The arXiv preprint is not yet available.

Authors: Hoigi Seo*, Byung Hyun Lee*, Minjun Kim*, Dohyun Mah, Jongho Lee and Se Young Chun. (* co-first author)