Skip to main navigation Skip to search Skip to main content

ERA: A QoE-Aware Collaborative Inference Algorithm for NOMA-Based Edge Intelligence

  • Xin Yuan
  • , Ning Li*
  • , Quan Chen
  • , Wenchao Xu
  • , Song Guo
  • *Corresponding author for this work
  • School of Ocean Engineering, Harbin Institute of Technology Weihai
  • School of Computer Science and Technology (School of Software), Harbin Institute of Technology Weihai
  • Guangdong University of Technology
  • Hong Kong University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Although AI has been extensively adopted and has profoundly transformed our lives, it is not feasible to directly deploy large AI models on edge devices with limited resources. To enhance the performance of Edge Intelligence (EI), model split inference has been proposed. In this approach, an AI model is segmented into sub-models, with the most resource-intensive parts offloaded wirelessly to the edge server. This reduces the resource demands and inference latency on the device. However, previous studies have primarily focused on enhancing and optimizing system Quality of Service (QoS), often overlooking Quality of Experience (QoE), which is another crucial aspect for users. Even though QoE has been extensively studied in Edge Computing (EC), the distinct differences between task offloading in EC and split inference in EI, along with specific QoE issues that remain unaddressed in both fields, render these algorithms ineffective for edge split inference scenarios. Therefore, this paper introduces an effective resource allocation algorithm, dubbed ERA, which aims to: 1) expedite split inference in EI, and 2) balance inference delay, QoE, and resource consumption. ERA incorporates resource consumption, QoE, and inference latency to determine the most optimal model split and resource allocation strategies. Given that it is impossible to simultaneously minimize inference delay and resource consumption while maximizing QoE, we employ a gradient descent-based algorithm to find the best possible compromise. Furthermore, to address the complexity arising from parameter discretization in the gradient descent algorithm, we have developed a pipeline gradient descent approach, known as PipGD. We have also examined the properties of the proposed algorithms, including their convergence, complexity, and approximation error. The experimental results clearly show that ERA outperforms previous studies significantly in terms of performance.

Original languageEnglish
Pages (from-to)2303-2319
Number of pages17
JournalIEEE Transactions on Mobile Computing
Volume25
Issue number2
DOIs
StatePublished - 2026
Externally publishedYes

Keywords

  • Edge intelligence
  • QoE
  • inference accelerating
  • model split

Fingerprint

Dive into the research topics of 'ERA: A QoE-Aware Collaborative Inference Algorithm for NOMA-Based Edge Intelligence'. Together they form a unique fingerprint.

Cite this