Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
We propose a phase diagram that determines whether alignment (CA) or prediction (CP) is effective—or neither—for multimodal learning, based on data characteristics.
Although CA and CP are widely used in multimodal representation learning, there is no systematic understanding of when each method succeeds or fails. In scientific domains such as biomedicine or astrophysics, heterogeneous instruments often cause standard methods to underperform the best single modality, yet there is no way to diagnose the cause.
We introduce a spiked signal-plus-noise model that explicitly models cross-modal correlation and noise structure. By analyzing the failure modes of CA and CP, we find that CA whitens each modality and fails when noise is strongly correlated across views, while CP's recovery performance is governed by source modality quality. Based on this, we construct a phase diagram with four regimes (Both, CA only, CP only, Neither) and propose a procedure to estimate a real dataset's location in this diagram using a small labeled subsample.
Experiments on synthetic data, stereo-vision benchmarks, image-caption pairs, and real astrophysical data validate the predictions even in nonlinear settings. Notably, we empirically demonstrate the existence of the Neither regime where cross-modal training is actively harmful, and provide a diagnostic tool to select the appropriate objective before training.