NeurIPS 2026✦ SpotlightCROSS-MODAL LEARNING, REIMAGINED

Strong helps
weak.

Directional Cross-Modal Alignment
Transfer in Multi-modal LLMs

A stronger vision model. A better audio or video model.
Transfer alignment across modalities—without further fine-tuning.

THE TRANSFER PRINCIPLECONCEPTUAL
◈
STRONG DONORVision MLLMWell-aligned with text
V
text-aligned directions
✳
DCATClosed-form alignment transfer
↓
Audio MLLMBetter listening
Video MLLMBetter understanding

Hoigi Seo1*Byung Hyun Lee1*Minjun Kim1*Dohyun Mah1Jongho Lee2Se Young Chun1,2†

1 Department of ECE   2 IPAI & INMC · Seoul National University

* Equal contribution   † Corresponding author

+29.55%

Audio performance

Reported average relative improvement
+24.92%

Video performance

Reported average relative improvement
256

Calibration samples

Per recipient modality
0

Gradient updates

A closed-form solution
01 / THE IDEA

Different modalities.
One transferable strength.

What if improving an audio model didn’t require more audio training?

Vision, audio, and video models can share the same language backbone, yet differ in how well their modality representations align with text. DCAT uses a well-aligned donor to improve a weaker recipient on its own tasks.

The key is directional alignment transfer: preserve the recipient’s capabilities while borrowing text-compatible directions from the donor.

Read the abstract +

Multi-modal large language models (MLLMs) achieve strong modality understanding by pairing a large language model (LLM) with an encoder for a target modality such as vision, video, or audio. However, improving an MLLM’s capability for a given modality typically requires additional training on large modality-specific datasets, incurring substantial data collection and compute costs. Model merging offers an alternative, but it is often infeasible for data-scarce, large per-sample size, or domain-specific modalities (e.g., audio and video), where same-modality model variants are rarely available. In this work, we characterize an intriguing asymmetric phenomenon: merging a well-aligned, data-rich source-modality MLLM into a data-scarce target-modality MLLM substantially improves the target on its own benchmarks. Our theoretical and empirical analyses show that this gain stems from enhanced alignment between modality-specific and textual tokens, induced by the stronger donor modality. Specifically, we derive a mutual-information lower bound that is monotonic in alignment-related quantities and strongly correlated with downstream MLLM performance. Building on this principle, we propose Directional Cross-Modal Alignment Transfer (DCAT), a novel framework that transfers textual alignment from a strong, well-aligned source (donor) modality to a weak target (recipient) modality, boosting target-modality performance without further fine-tuning. We further show that the alignment-enhancing objective admits a closed-form weight-space solution computed from only a small calibration set. DCAT outperforms existing model-merging methods, offering an efficient path toward cross-modal alignment transfer.

02 / HOW IT WORKS

Alignment is the bridge.

Transfer the right directions.
Preserve what the recipient already knows.

01

Find the shared subspace

Use a small calibration set to identify the principal text subspace in the recipient’s activations.

Text + modality activations
02

Build a better target

Project modality activations into the text subspace and balance their spectrum while preserving total energy.

Overlap + spectral diversity
03

Transfer, layer by layer

Solve a quadratic objective in closed form, guided by donor and recipient anchors. Recompute activations for the next layer.

No gradient-based fine-tuning
THE ALIGNMENT SCORE
Alignγ = SEOγ · SD1−γ

Subspace Energy Overlap measures how much modality energy enters the text subspace.
Spectral Diversity measures how broadly it spreads across that subspace.

The score motivates a mutual-information lower bound under the assumptions in Theorem 1. It is an alignment proxy, not a direct measurement of mutual information.

03 / THE EVIDENCE

Small calibration.
Meaningful gains.

Explore the reported benchmark scores.
Same recipient. Stronger alignment.

Vision → Audio

Vicuna-based models · Higher is better

OriginalDCAT
Scores transcribed from Table 2. Relative gains are calculated from the displayed, rounded scores. Bar lengths use a common 0–100 scale.
Compare all merging methods +

Source: Table 2 in the paper. The headline improvements above follow the paper’s reported summaries; the chart displays individual benchmark scores.

04 / WHY ALIGNMENT MATTERS

A mechanism backed by evidence.

114checkpoints

Alignment tracks performance.

Across merged Vicuna checkpoints, the proposed score shows strong positive correlations with downstream performance.

SEO + SDincrease after DCAT

Better alignment. Preserved scale.

DCAT increases overlap and spectral diversity while keeping the modality-token norm ratio close to one.

Scope: the main experiments use vision donors and audio/video recipients with a shared LLM architecture. The direction matters; these results do not imply that every modality pair benefits equally.

05 / BUILD ON THIS WORK

Cite DCAT.

Strong helps weak.
Directional Cross-Modal Alignment Transfer.

BIBTEX
@inproceedings{seo2026strong,
  title={Strong Helps Weak: Directional
    Cross-Modal Alignment Transfer
    in Multi-modal LLMs},
  author={Seo, Hoigi and Lee, Byung Hyun
    and Kim, Minjun and Mah, Dohyun
    and Lee, Jongho and Chun, Se Young},
  booktitle={Advances in Neural Information
    Processing Systems},
  year={2026}
}

Open full-resolution figure ↗