Skip to main navigation Skip to search Skip to main content

Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference

  • Cheng Yuan
  • , Zhening Liu
  • , Jiashu Lv
  • , Jiawei Shao
  • , Yufei Jiang
  • , Jun Zhang
  • , Xuelong Li*
  • *Corresponding author for this work
  • China Telecommunications
  • Harbin Institute of Technology Shenzhen
  • Hong Kong University of Science and Technology
  • Peking University

Research output: Contribution to journalArticlepeer-review

Abstract

With the rapid development of large multimodal models (LMMs), multimodal understanding applications are emerging. As most LMM inference requests originate from edge devices with limited computational capabilities, the predominant inference pipeline involves directly forwarding the input data to an edge server which handles all computations. However, this approach introduces high transmission latency due to limited uplink bandwidth of edge devices and significant computation latency caused by the prohibitive number of visual tokens, thus hindering delay-sensitive tasks and degrading user experience. To address this challenge, we propose a task-oriented feature compression (TOFC) method for multimodal understanding in a device-edge co-inference framework, where visual features are merged by clustering and encoded by a learnable and selective entropy model before feature projection. Specifically, we employ density peaks clustering based on KK nearest neighbors to reduce the number of visual features, thereby minimizing both data transmission and computational complexity. Subsequently, a learnable entropy model with hyperprior is utilized to encode and decode merged features, further reducing transmission overhead. To enhance compression efficiency, multiple entropy models are adaptively selected based on the characteristics of the visual features, enabling a more accurate estimation of the probability distribution. Comprehensive experiments on seven visual question answering benchmarks validate the effectiveness of the proposed TOFC method. Results show that TOFC achieves up to 52% reduction in data transmission overhead and 63% reduction in system latency while maintaining identical task performance, compared with neural compression ELIC.

Original languageEnglish
Pages (from-to)4762-4775
Number of pages14
JournalIEEE Transactions on Mobile Computing
Volume25
Issue number4
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Distributed inference
  • edge artificial intelligence
  • large multimodal models
  • task-oriented communication

Fingerprint

Dive into the research topics of 'Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference'. Together they form a unique fingerprint.

Cite this