Skip to main navigation Skip to search Skip to main content

BiTA: Bi-directional tuning for lossless acceleration in large language models

  • Feng Lin*
  • , Hanling Yi
  • , Yifan Yang
  • , Hongbin Li
  • , Xiaotian Yu
  • , Guangming Lu
  • , Rong Xiao
  • *Corresponding author for this work
  • Intellifusion
  • Harbin Institute of Technology Shenzhen

Research output: Contribution to journalArticlepeer-review

Abstract

Large language models (LLMs) typically employ autoregressive generation during inference, leading to high memory bandwidth demand and consequently extended latency. An effective strategy to mitigate this inefficiency is speculative decoding, which reduces the number of model inference calls, thereby lowering memory bandwidth requirements. In this paper, we propose BiTA (Bi-directional Tuning for lossless Acceleration), an innovative speculative decoding method that expedites LLMs through streamlined semi-autoregressive generation and draft verification. BiTA enhances LLMs with a parameter-efficient design called bi-directional tuning, enabling semi-autoregressive generation, while leveraging an efficient tree-based decoding mechanism to perform draft candidate generation and verification in parallel, ensuring that the outputs of accelerated LLMs remain identical to those of their original autoregressive counterparts. As a lightweight plug-in module, BiTA seamlessly boosts the inference efficiency of existing LLMs without requiring additional assistance models or incurring significant extra memory costs. Applying BiTA, LLaMA-2-70B-Chat achieves a 2.7× speedup on the MT-Bench benchmark. Extensive experiments confirm that BiTA surpasses state-of-the-art speculative decoding methods. The code is available at https://github.com/linfeng93/BiTA.

Original languageEnglish
Article number127305
JournalExpert Systems with Applications
Volume279
DOIs
StatePublished - 15 Jun 2025
Externally publishedYes

Keywords

  • Draft verification
  • LLM
  • Lossless acceleration
  • Prompt tuning
  • Semi-autoregressive generation
  • Speculative decoding

Fingerprint

Dive into the research topics of 'BiTA: Bi-directional tuning for lossless acceleration in large language models'. Together they form a unique fingerprint.

Cite this