Abstract
Large Vision-Language Models (LVLMs) demonstrate strong performance on high-level semantic tasks but lack fine-grained perception for pixel-level understanding, particularly in Referring Expression Segmentation (RES). The core challenge is representing irregular object contours compatibly with the sequential nature of language models. We introduce Seg-LLaVA, an end-to-end LVLM that redefines segmentation by predicting contour points in polar coordinates rather than dense pixel masks. Our Polar Coordinate Adaptive Sampling (PCAS) samples key boundary points, providing a unified representation that enhances training stability and shape fidelity. A lightweight refinement module leverages hidden states to generate high-precision masks with improved boundaries. We also introduce LAGS, a large-scale dataset enabling complex, interactive, and multi-object video segmentation. Extensive experiments show Seg-LLaVA achieves state-of-the-art performance, substantially surpassing previous methods in segmentation accuracy, localization precision, and language grounding.
| Original language | English |
|---|---|
| Article number | 113560 |
| Journal | Pattern Recognition |
| Volume | 179 |
| DOIs | |
| State | Published - Nov 2026 |
| Externally published | Yes |
Keywords
- Large language models
- Polar coordinate adaptive
- Referring expression segmentation
- Segmentation
Fingerprint
Dive into the research topics of 'Seg-LLaVA: Empowering pixel-level understanding with large vision language model'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver