TY - GEN
T1 - Growing a Twig to Accelerate Large Vision-Language Models
AU - Shao, Zhenwei
AU - Wang, Mingyang
AU - Yu, Zhou
AU - Pan, Wenwen
AU - Yang, Yan
AU - Wei, Tao
AU - Zhang, Hongyuan
AU - Mao, Ning
AU - Chen, Wei
AU - Yu, Jun
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens guided by the attention maps of VLM's early layers. Despite the success of these token pruning methods, they still suffer from two major shortcomings: (i) considerable accuracy drop due to insensitive attention signals in early layers, and (ii) limited speedup when generating long responses (e.g., 30 tokens). To address the limitations above, we present TwigVLM-a simple and general architecture by 'growing' a lightweight twig upon an early layer of the base VLM. Compared with most existing VLM acceleration methods purely based on visual token pruning, our TwigVLM not only achieves better accuracy retention by employing a twig-guided token pruning (TTP) strategy, but also yields higher generation speed by utilizing a self-speculative decoding (SSD) strategy. Taking LLaVA-1.5-7B as the base VLM, experimental results show that TwigVLM preserves 96% of the original performance after pruning 88.9% of visual tokens and achieves 154% speedup in generating long responses, delivering significantly better performance in terms of both accuracy and speed over the state-of-the-art VLM acceleration methods.
AB - Large vision-language models (VLMs) have demonstrated remarkable capabilities in open-world multimodal understanding, yet their high computational overheads pose great challenges for practical deployment. Some recent works have proposed methods to accelerate VLMs by pruning redundant visual tokens guided by the attention maps of VLM's early layers. Despite the success of these token pruning methods, they still suffer from two major shortcomings: (i) considerable accuracy drop due to insensitive attention signals in early layers, and (ii) limited speedup when generating long responses (e.g., 30 tokens). To address the limitations above, we present TwigVLM-a simple and general architecture by 'growing' a lightweight twig upon an early layer of the base VLM. Compared with most existing VLM acceleration methods purely based on visual token pruning, our TwigVLM not only achieves better accuracy retention by employing a twig-guided token pruning (TTP) strategy, but also yields higher generation speed by utilizing a self-speculative decoding (SSD) strategy. Taking LLaVA-1.5-7B as the base VLM, experimental results show that TwigVLM preserves 96% of the original performance after pruning 88.9% of visual tokens and achieves 154% speedup in generating long responses, delivering significantly better performance in terms of both accuracy and speed over the state-of-the-art VLM acceleration methods.
KW - inference acceralation
KW - large vision-language model
KW - self speculative decoding
KW - visual token pruning
UR - https://www.scopus.com/pages/publications/105044137317
U2 - 10.1109/ICCV51701.2025.01866
DO - 10.1109/ICCV51701.2025.01866
M3 - 会议稿件
AN - SCOPUS:105044137317
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 20064
EP - 20074
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Y2 - 19 October 2025 through 23 October 2025
ER -