Abstract
Recently, vision transformer (ViT) has demonstrated remarkable efficacy across multiple visual tasks. However, prevalent paradigms of ViT variants involve the optimization of a singular model for a particular task, resulting in a linear escalation in total model size as the number of tasks increases. To address this issue, we develop Poly-TF, a polymeric transformer framework tailored to concurrently optimize multiple visual tasks spanning diverse dataset domains. The core of Poly-TF is leveraging shared parameters to reduce total storage requisites, along with the utilization of task-specific parameters to capture distinctive feature representations for each task. Technically, we develop a prompt-based modulation mechanism that seeks to compensate discrepancies across diverse task domains, thereby empowering the network to gain a deep understanding for each prevailing task and facilitating task-specific feature extraction. Besides, we introduce an adaptive granular parameter-sharing scheme applied to each linear layer within the transformer block, which automatically discerns the shared parameters for efficient storage and task-specific parameters for further specific task perception. Benefitting from the proposed techniques, Poly-TF explicitly establishes feature representations unique to each task via the task-specific components while implicitly modeling generic information along with cross-task correlations through the shared ones. Rigorous experiments prove that Poly-TF exhibits competitive performance in various visual tasks with substantially reduced storage demands. Moreover, we also demonstrate that the presence of task-specific components in Poly-TF intrinsically facilitates incremental learning while circumventing catastrophic forgetting. The code is available at https://github.com/Mrfanxuan/poly-tf.
| Original language | English |
|---|---|
| Article number | 114167 |
| Journal | Knowledge-Based Systems |
| Volume | 327 |
| DOIs | |
| State | Published - 9 Oct 2025 |
Keywords
- Multi-task learning
- Prompting
- Vision transformer
Fingerprint
Dive into the research topics of 'Poly-TF: A polymeric transformer framework for multiple visual tasks at once'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver