Skip to main navigation Skip to search Skip to main content

Prior-driven multi-view clustering: a perspective from geometry and semantics to LMM cognition

  • Xuqian Xue
  • , Jie Wen
  • , Xinwang Liu
  • , Junping Zhang*
  • *Corresponding author for this work
  • Fudan University
  • Harbin Institute of Technology
  • National University of Defense Technology

Research output: Contribution to journalArticlepeer-review

Abstract

With the rapid emergence of large multimodal models (LMMs), massive heterogeneous data are proliferating across various industrial and scientific domains. In this context, multi-view clustering (MVC) serves as a cornerstone technology for unsupervised knowledge discovery and latent correlation mining. At present, MVC is undergoing a profound and historical paradigm shift. Traditional surveys predominantly focus on the horizontal categorization of algorithmic network structures. However, this approach often fails to reveal the intrinsic evolutionary logic across different technological eras. Departing from these conventions, this study proposes a pioneering, prior-driven theoretical perspective to reconstruct systematically the developmental trajectory of MVC over the past two decades. This goal is achieved through a trans-paradigm analytical framework of geometry-semantics-cognition. During the initial stage of shallow structural mining, the research paradigm focused on explicit mathematical constraints within original or kernel-induced feature spaces. Euclidean space methods, such as multi-view K-Means and nonnegative matrix factorization, identify global prototypes by minimizing squared error or Frobenius norm reconstruction loss. By contrast, affine space methods leverage self-representation properties to model data as a union of low-dimensional subspaces, while manifold space techniques utilize spectral graph theory to transform clustering into optimal graph cut problems by preserving local topological correlations. As the field transitioned into deep spatial modeling based on semantic collaborative priors, researchers utilized the powerful nonlinear mapping capabilities of deep neural networks to project heterogeneous data into high-order semantic spaces. This evolution encompasses several distinct research paradigms. 1) Embedding space research focuses on deep subspace clustering, employing autoencoders to learn discriminative features while maintaining cross-view consistency. 2) Latent space methods utilize probabilistic generative models, such as variational autoencoders and generative adversarial networks, to align latent distributions and infer missing view information through adversarial games. 3) Augmented space paradigms introduce contrastive learning to maximize mutual information between views via InfoNCE-like losses, enhancing representation robustness. 4) Topological space studies leverage graph neural networks to mine intra-view geometric structures and inter-view complementary semantics synergistically. Moving into the current era of LMMs, this study progressively explores deep alignment based on cognitive priors, where LMMs are viewed not merely as data sources but as knowledge bases that contain human-level common sense and logical reasoning capability. We systematically elucidate the infrastructural role of MVC in empowering massive data governance. Specifically, MVC facilitates semantic dedduplication to enhance data quality and employs token-level clustering to optimize mixture-of-experts (MoE) routing for expert specialization. Furthermore, MVC enables hierarchical semantic chunking, which is critical for precise document retrieval within retrieval-augmented generation (RAG) frameworks. Beyond these applications, we analyze the potential for LMM logical reasoning to back-propagate into clustering tasks. This synergy elevates MVC from pure statistical feature alignment to a new dimension of knowledge-driven cognitive logic consistency. Beyond theoretical frameworks, the review highlights the transformative effect of MVC in diverse real-world scenarios, ranging from multi-omics fusion in smart healthcare for cancer subtyping and spatial-spectral fusion in remote sensing for urban functional zone identification to cross-perspective pedestrian recognition in public security and unsupervised anomaly detection in industrial internet of things networks. Despite these advancements, achieving a transition from laboratory benchmarks to robust industrial infrastructure requires addressing several core challenges. First, the scaling bottleneck where in the O(N2) or O(N3) computational complexity of traditional methods must be reduced to linear levels to handle million-scale web data. Second, the extreme robustness required in open environments to handle not only severely incomplete views but also the pervasive long-tailed category imbalance where minority abnormal samples are frequently overwhelmed by dominant patterns. Third, the urgent need for evaluation system reconstruction to move away from legacy small-scale datasets with shallow features toward native heterogeneous multimodal benchmarks that reflect real-world weak alignment and complex noise distributions. Ultimately, this survey aims to provide a novel research road map for theoretical innovation and engineering practice, advocating for a transition toward intelligent decision systems characterized by high-level semantic decoupling and cognitive logic alignment in the era of LMMs.

Translated title of the contribution先验驱动的多视图聚类:从几何与语义空间到多模态大模型认知前瞻
Original languageEnglish
Pages (from-to)2408-2433
Number of pages26
JournalJournal of Image and Graphics
Volume31
Issue number7
DOIs
StatePublished - 2026
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being
  2. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure
  3. SDG 11 - Sustainable Cities and Communities
    SDG 11 Sustainable Cities and Communities

Keywords

  • cognitive alignment
  • geometric structure
  • large multimodal models (LMMs)
  • multi-view clustering (MVC)
  • prior-driven learning
  • review
  • semantic collaboration

Fingerprint

Dive into the research topics of 'Prior-driven multi-view clustering: a perspective from geometry and semantics to LMM cognition'. Together they form a unique fingerprint.

Cite this