Skip to main navigation Skip to search Skip to main content

Growing a Twig to Accelerate Large Vision-Language Models

  • Zhenwei Shao
  • , Mingyang Wang
  • , Zhou Yu*
  • , Wenwen Pan
  • , Yan Yang
  • , Tao Wei
  • , Hongyuan Zhang
  • , Ning Mao
  • , Wei Chen
  • , Jun Yu
  • *Corresponding author for this work
  • Hangzhou Dianzi University
  • Li Auto Inc.
  • Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens guided by the attention maps of VLM's early layers. Despite the success of these token pruning methods, they still suffer from two major shortcomings: (i) considerable accuracy drop due to insensitive attention signals in early layers, and (ii) limited speedup when generating long responses (e.g., 30 tokens). To address the limitations above, we present TwigVLM-a simple and general architecture by 'growing' a lightweight twig upon an early layer of the base VLM. Compared with most existing VLM acceleration methods purely based on visual token pruning, our TwigVLM not only achieves better accuracy retention by employing a twig-guided token pruning (TTP) strategy, but also yields higher generation speed by utilizing a self-speculative decoding (SSD) strategy. Taking LLaVA-1.5-7B as the base VLM, experimental results show that TwigVLM preserves 96% of the original performance after pruning 88.9% of visual tokens and achieves 154% speedup in generating long responses, delivering significantly better performance in terms of both accuracy and speed over the state-of-the-art VLM acceleration methods.

Original languageEnglish
Title of host publicationProceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages20064-20074
Number of pages11
ISBN (Electronic)9798331587758
DOIs
StatePublished - 2025
Externally publishedYes
Event2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, United States
Duration: 19 Oct 202523 Oct 2025

Publication series

NameProceedings of the IEEE International Conference on Computer Vision
ISSN (Print)1550-5499
ISSN (Electronic)2380-7504

Conference

Conference2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Country/TerritoryUnited States
CityHonolulu
Period19/10/2523/10/25

Keywords

  • inference acceralation
  • large vision-language model
  • self speculative decoding
  • visual token pruning

Fingerprint

Dive into the research topics of 'Growing a Twig to Accelerate Large Vision-Language Models'. Together they form a unique fingerprint.

Cite this