Skip to main navigation Skip to search Skip to main content

Poly-TF: A polymeric transformer framework for multiple visual tasks at once

  • Harbin Institute of Technology
  • State Key Yangtze River Delta HIT Robot Technology Research Institute

Research output: Contribution to journalArticlepeer-review

Abstract

Recently, vision transformer (ViT) has demonstrated remarkable efficacy across multiple visual tasks. However, prevalent paradigms of ViT variants involve the optimization of a singular model for a particular task, resulting in a linear escalation in total model size as the number of tasks increases. To address this issue, we develop Poly-TF, a polymeric transformer framework tailored to concurrently optimize multiple visual tasks spanning diverse dataset domains. The core of Poly-TF is leveraging shared parameters to reduce total storage requisites, along with the utilization of task-specific parameters to capture distinctive feature representations for each task. Technically, we develop a prompt-based modulation mechanism that seeks to compensate discrepancies across diverse task domains, thereby empowering the network to gain a deep understanding for each prevailing task and facilitating task-specific feature extraction. Besides, we introduce an adaptive granular parameter-sharing scheme applied to each linear layer within the transformer block, which automatically discerns the shared parameters for efficient storage and task-specific parameters for further specific task perception. Benefitting from the proposed techniques, Poly-TF explicitly establishes feature representations unique to each task via the task-specific components while implicitly modeling generic information along with cross-task correlations through the shared ones. Rigorous experiments prove that Poly-TF exhibits competitive performance in various visual tasks with substantially reduced storage demands. Moreover, we also demonstrate that the presence of task-specific components in Poly-TF intrinsically facilitates incremental learning while circumventing catastrophic forgetting. The code is available at https://github.com/Mrfanxuan/poly-tf.

Original languageEnglish
Article number114167
JournalKnowledge-Based Systems
Volume327
DOIs
StatePublished - 9 Oct 2025

Keywords

  • Multi-task learning
  • Prompting
  • Vision transformer

Fingerprint

Dive into the research topics of 'Poly-TF: A polymeric transformer framework for multiple visual tasks at once'. Together they form a unique fingerprint.

Cite this