Skip to main navigation Skip to search Skip to main content

CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks

  • Mingshuang Luo
  • , Ruibing Hou*
  • , Bo Chao
  • , Hong Chang
  • , Zimo Liu
  • , Yaowei Wang
  • , Shiguang Shan
  • *Corresponding author for this work
  • CAS - Institute of Computing Technology
  • University of Chinese Academy of Sciences
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

Human-centric visual analysis plays a pivotal role in diverse applications, including surveillance, healthcare, and human-computer interaction. With the emergence of large-scale unlabeled human image datasets, there is an increasing need for a general unsupervised pre-training model capable of supporting diverse human-centric downstream tasks. To achieve this goal, we propose CLASP (CLIP-guided Adaptable Self-suPervised learning), a novel framework designed for unsupervised pre-training in human-centric visual tasks. CLASP leverages the powerful vision-language model CLIP to generate both low-level (e.g. body parts) and high-level (e.g. attributes) semantic pseudo-labels. These multi-level semantic cues are then integrated into the learned visual representations, enriching their expressiveness and generalizability. Recognizing that different downstream tasks demand varying levels of semantic granularity, CLASP incorporates a Prompt-Controlled Mixture-of-Experts (MoE) module. MoE dynamically adapts feature extraction based on task-specific prompts, mitigating potential feature conflicts and enhancing transferability. Furthermore, CLASP employs a multi-task pre-training strategy, where part- and attribute-level pseudo-labels derived from CLIP guide the representation learning process. Extensive experiments across multiple benchmarks demonstrate that CLASP consistently outperforms existing unsupervised pre-training methods, advancing the field of human-centric visual analysis.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 3 - Good Health and Well-being
    SDG 3 Good Health and Well-being

Keywords

  • human-centric visual tasks
  • mixture-of-experts
  • self-supervised learning

Fingerprint

Dive into the research topics of 'CLIP-Guided Adaptable Self-Supervised Learning for Human-Centric Visual Tasks'. Together they form a unique fingerprint.

Cite this