蔡逸超
我目前是阿德莱德大学, 澳大利亚机器学习研究所 (AIML) 的计算机科学博士生,导师是 史勤峰 教授。我本科和硕士毕业于武汉理工大学,并曾在加州大学伯克利分校,California PATH 做过访问学生研究员。
我的研究关注:在什么条件下,学习目标能够以可证明的方式识别出支持泛化的因果潜在结构,以及它们在什么情况下无法做到这一点。在方法上,我主要结合可识别性理论、潜变量建模、总体目标函数分析以及表征几何等工具来研究这一问题。在实证方面,我主要研究视觉—语言模型和自回归模型,并逐步拓展至生成建模与科学表征学习。(浏览交互式研究议程 →)
代表性论文
-
The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence
International Conference on Machine Learning (ICML), 2026
Abstract
While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment–uniformity decomposition. We develop a measure-theoretic framework in which representation measures evolve on a fixed embedding manifold. In the large-batch limit, we prove value and gradient consistency, linking the stochastic objective to explicit deterministic energy landscapes and revealing a geometric bifurcation between unimodal and symmetric multimodal regimes. In the unimodal case, the intrinsic energy is strictly convex and admits a unique Gibbs equilibrium, showing that entropy acts as a tie-breaker within the aligned basin. In the multimodal case, the intrinsic geometry becomes cross-coupled and contains a persistent negative symmetric divergence term: each modality’s marginal reshapes the effective landscape of the other, allowing strong pairwise alignment to coexist with a persistent modality gap. Controlled synthetic experiments and analyses of pretrained CLIP representations support these predictions. Overall, our results shift the analytical lens from pointwise discrimination to population geometry, showing that pairwise alignment alone is insufficient to control cross-modal marginal structure.
BibTeX
@inproceedings{cai2026geometric, title = {The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence}, author = {Cai, Yichao and Zhang, Zhen and Liu, Yuhang and Shi, Javen Q.}, booktitle = {International Conference on Machine Learning (ICML)}, year = {2026}, } -
On the Value of Cross-Modal Misalignment in Multimodal Representation Learning
Advances in Neural Information Processing Systems (NeurIPS), 2025 · Spotlight
Abstract
Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered – both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning.
BibTeX
@inproceedings{cai2025misalignment, title = {On the Value of Cross-Modal Misalignment in Multimodal Representation Learning}, author = {Cai, Yichao and Liu, Yuhang and Gao, Erdun and Jiang, Tianjiao and Zhang, Zhen and van den Hengel, Anton and Shi, Javen Q.}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2025}, }
最新文章
What Does InfoNCE Actually Do to Representation Geometry?
How InfoNCE shapes representation distributions through alignment potentials, entropic dispersion, and cross-modal coupling.