Find the shared subspace
Use a small calibration set to identify the principal text subspace in the recipient’s activations.
Text + modality activationsA stronger vision model. A better audio or video model.
Transfer alignment across modalities—without further fine-tuning.
Audio performance
Reported average relative improvementVideo performance
Reported average relative improvementCalibration samples
Per recipient modalityGradient updates
A closed-form solutionWhat if improving an audio model didn’t require more audio training?
Vision, audio, and video models can share the same language backbone, yet differ in how well their modality representations align with text. DCAT uses a well-aligned donor to improve a weaker recipient on its own tasks.
The key is directional alignment transfer: preserve the recipient’s capabilities while borrowing text-compatible directions from the donor.
Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM’s capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (e.g., audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-Modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer.
Transfer the right directions.
Preserve what the recipient already knows.
Use a small calibration set to identify the principal text subspace in the recipient’s activations.
Text + modality activationsProject modality activations into the text subspace and balance their spectrum while preserving total energy.
Overlap + spectral diversitySolve a quadratic objective in closed form, guided by donor and recipient anchors. Recompute activations for the next layer.
No gradient-based fine-tuningSubspace Energy Overlap measures how much modality energy enters the text subspace.
Spectral Diversity measures how broadly it spreads across that subspace.
The score motivates a mutual-information lower bound under the assumptions in Theorem 1. It is an alignment proxy, not a direct measurement of mutual information.
Explore the reported benchmark scores.
Same recipient. Stronger alignment.
Vicuna-based models · Higher is better
Source: Table 2 in the paper. The headline improvements above follow the paper’s reported summaries; the chart displays individual benchmark scores.
Across merged Vicuna checkpoints, the proposed score shows strong positive correlations with downstream performance.
DCAT increases overlap and spectral diversity while keeping the modality-token norm ratio close to one.
Scope: the main experiments use vision donors and audio/video recipients with a shared LLM architecture. The direction matters; these results do not imply that every modality pair benefits equally.
Strong helps weak.
Directional Cross-Modal Alignment Transfer.
@inproceedings{seo2026strong,
title={Strong Helps Weak: Directional
Cross-Modal Alignment Transfer
in Multi-modal LLMs},
author={Seo, Hoigi and Lee, Byung Hyun
and Kim, Minjun and Mah, Dohyun
and Lee, Jongho and Chun, Se Young},
booktitle={Advances in Neural Information
Processing Systems},
year={2026}
}