Yichao Cai

I am a final-year PhD student in Computer Science at the Australian Institute for Machine Learning (AIML), Adelaide University, advised by Prof. Javen Qinfeng Shi. I received my M.Sc. and B.Eng. degrees from Wuhan University of Technology and spent time as a visiting student researcher at California PATH, UC Berkeley.

My research asks when learning objectives provably identify latent structure that supports generalization, and when they fail to do so. I approach this question using identifiability theory, latent-variable modeling, population-level objective analysis, and representation geometry. Empirically, I study vision-language and autoregressive models, with growing interests in generative modeling and scientific representation learning. (Explore my research agenda →)

Selected publications

  1. The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence

    Yichao Cai, Zhen Zhang, Yuhang Liu, and Javen Q. Shi

    International Conference on Machine Learning (ICML), 2026

    Conference Contrastive learningRepresentation geometryCross-modal alignmentVision-language
    Abstract

    While InfoNCE underlies modern contrastive learning, its geometric mechanisms remain under-characterized beyond the canonical alignment–uniformity decomposition. We develop a measure-theoretic framework in which representation measures evolve on a fixed embedding manifold. In the large-batch limit, we prove value and gradient consistency, linking the stochastic objective to explicit deterministic energy landscapes and revealing a geometric bifurcation between unimodal and symmetric multimodal regimes. In the unimodal case, the intrinsic energy is strictly convex and admits a unique Gibbs equilibrium, showing that entropy acts as a tie-breaker within the aligned basin. In the multimodal case, the intrinsic geometry becomes cross-coupled and contains a persistent negative symmetric divergence term: each modality’s marginal reshapes the effective landscape of the other, allowing strong pairwise alignment to coexist with a persistent modality gap. Controlled synthetic experiments and analyses of pretrained CLIP representations support these predictions. Overall, our results shift the analytical lens from pointwise discrimination to population geometry, showing that pairwise alignment alone is insufficient to control cross-modal marginal structure.

    BibTeX
    @inproceedings{cai2026geometric,
      title = {The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-Modal Divergence},
      author = {Cai, Yichao and Zhang, Zhen and Liu, Yuhang and Shi, Javen Q.},
      booktitle = {International Conference on Machine Learning (ICML)},
      year = {2026},
    }
    
  2. On the Value of Cross-Modal Misalignment in Multimodal Representation Learning

    Yichao Cai, Yuhang Liu, Erdun Gao, Tianjiao Jiang, Zhen Zhang, Anton van den Hengel, and Javen Q. Shi

    Advances in Neural Information Processing Systems (NeurIPS), 2025 · Spotlight

    Conference Contrastive learningIdentifiabilityCross-modal alignmentVision-language
    Abstract

    Multimodal representation learning, exemplified by multimodal contrastive learning (MMCL) using image-text pairs, aims to learn powerful representations by aligning cues across modalities. This approach relies on the core assumption that the exemplar image-text pairs constitute two representations of an identical concept. However, recent research has revealed that real-world datasets often exhibit cross-modal misalignment. There are two distinct viewpoints on how to address this issue: one suggests mitigating the misalignment, and the other leveraging it. We seek here to reconcile these seemingly opposing perspectives, and to provide a practical guide for practitioners. Using latent variable models we thus formalize cross-modal misalignment by introducing two specific mechanisms: Selection bias, where some semantic variables are absent in the text, and perturbation bias, where semantic variables are altered – both leading to misalignment in data pairs. Our theoretical analysis demonstrates that, under mild assumptions, the representations learned by MMCL capture exactly the information related to the subset of the semantic variables invariant to selection and perturbation biases. This provides a unified perspective for understanding misalignment. Based on this, we further offer actionable insights into how misalignment should inform the design of real-world ML systems. We validate our theoretical findings via extensive empirical studies on both synthetic data and real image-text datasets, shedding light on the nuanced impact of cross-modal misalignment on multimodal representation learning.

    * Yichao Cai and Yuhang Liu contributed equally.

    BibTeX
    @inproceedings{cai2025misalignment,
      title = {On the Value of Cross-Modal Misalignment in Multimodal Representation Learning},
      author = {Cai, Yichao and Liu, Yuhang and Gao, Erdun and Jiang, Tianjiao and Zhang, Zhen and van den Hengel, Anton and Shi, Javen Q.},
      booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
      year = {2025},
    }
    

All publications →

Latest writing

What Does InfoNCE Actually Do to Representation Geometry?

How InfoNCE shapes representation distributions through alignment potentials, entropic dispersion, and cross-modal coupling.