Skip to main navigation Skip to search Skip to main content

Seg-LLaVA: Empowering pixel-level understanding with large vision language model

  • Fan Yang
  • , Yousong Zhu
  • , Yufei Zhan
  • , Hongyin Zhao
  • , Xin Li
  • , Yaowei Wang
  • , Ming Tang
  • , Xin Ning*
  • , Jinqiao Wang
  • *Corresponding author for this work
  • CAS - Institute of Automation
  • University of Chinese Academy of Sciences
  • Peng Cheng Laboratory
  • CAS - Institute of Semiconductors
  • Wuhan AI Research

Research output: Contribution to journalArticlepeer-review

Abstract

Large Vision-Language Models (LVLMs) demonstrate strong performance on high-level semantic tasks but lack fine-grained perception for pixel-level understanding, particularly in Referring Expression Segmentation (RES). The core challenge is representing irregular object contours compatibly with the sequential nature of language models. We introduce Seg-LLaVA, an end-to-end LVLM that redefines segmentation by predicting contour points in polar coordinates rather than dense pixel masks. Our Polar Coordinate Adaptive Sampling (PCAS) samples key boundary points, providing a unified representation that enhances training stability and shape fidelity. A lightweight refinement module leverages hidden states to generate high-precision masks with improved boundaries. We also introduce LAGS, a large-scale dataset enabling complex, interactive, and multi-object video segmentation. Extensive experiments show Seg-LLaVA achieves state-of-the-art performance, substantially surpassing previous methods in segmentation accuracy, localization precision, and language grounding.

Original languageEnglish
Article number113560
JournalPattern Recognition
Volume179
DOIs
StatePublished - Nov 2026
Externally publishedYes

Keywords

  • Large language models
  • Polar coordinate adaptive
  • Referring expression segmentation
  • Segmentation

Fingerprint

Dive into the research topics of 'Seg-LLaVA: Empowering pixel-level understanding with large vision language model'. Together they form a unique fingerprint.

Cite this