Skip to main navigation Skip to search Skip to main content

ServerlessLego: An Elastic Serverless Framework Assembling Model Building Blocks to Provide SLO-Aware Inference Services

  • Harbin Institute of Technology
  • Harbin Institute of Technology
  • Pengcheng Laboratory

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Inference of large language models (LLMs) is common in cloud environments. As the elastic resource management capabilities and the flexible pay-as-you-go billing model offered by serverless, LLM inference services are increasingly migrated to serverless platforms. However, the increasing size of LLMs in recent years has introduced a new cold start issue for serverless frameworks, which in turn impacts their scalability under dynamic workloads. To address these issues, we propose ServerlessLego, an elastic serverless computing framework. ServerlessLego partitions LLMs into layers, then groups and deploys them to different instances, and loads these groups in parallel. These instances perform a subscription-based pipeline. To address dynamically request loads, ServerlessLego models the incoming request patterns and the inference time of running requests, providing an SLO-Aware instance scheduling. Experiments show that ServerlessLego reduces the cold start time of serverless frameworks by 58.15 % and improves throughput by 43.39 % compared to the baseline for dynamic workloads. Moreover, ServerlessLego can horizontally schedule instance based on request SLOs and arrival rates.

Original languageEnglish
Title of host publicationProceedings of 2025 IEEE 31st International Conference on Parallel and Distributed Systems, ICPADS 2025
PublisherIEEE Computer Society
ISBN (Electronic)9798331549015
DOIs
StatePublished - 2025
Externally publishedYes
Event31st IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025 - Hefei, China
Duration: 14 Dec 202517 Dec 2025

Publication series

NameProceedings of the International Conference on Parallel and Distributed Systems - ICPADS
ISSN (Print)1521-9097

Conference

Conference31st IEEE International Conference on Parallel and Distributed Systems, ICPADS 2025
Country/TerritoryChina
CityHefei
Period14/12/2517/12/25

Keywords

  • Cold Start
  • LLM inference
  • Scalability
  • Serverless

Fingerprint

Dive into the research topics of 'ServerlessLego: An Elastic Serverless Framework Assembling Model Building Blocks to Provide SLO-Aware Inference Services'. Together they form a unique fingerprint.

Cite this