Abstract
Large language models (LLMs) typically employ autoregressive generation during inference, leading to high memory bandwidth demand and consequently extended latency. An effective strategy to mitigate this inefficiency is speculative decoding, which reduces the number of model inference calls, thereby lowering memory bandwidth requirements. In this paper, we propose BiTA (Bi-directional Tuning for lossless Acceleration), an innovative speculative decoding method that expedites LLMs through streamlined semi-autoregressive generation and draft verification. BiTA enhances LLMs with a parameter-efficient design called bi-directional tuning, enabling semi-autoregressive generation, while leveraging an efficient tree-based decoding mechanism to perform draft candidate generation and verification in parallel, ensuring that the outputs of accelerated LLMs remain identical to those of their original autoregressive counterparts. As a lightweight plug-in module, BiTA seamlessly boosts the inference efficiency of existing LLMs without requiring additional assistance models or incurring significant extra memory costs. Applying BiTA, LLaMA-2-70B-Chat achieves a 2.7× speedup on the MT-Bench benchmark. Extensive experiments confirm that BiTA surpasses state-of-the-art speculative decoding methods. The code is available at https://github.com/linfeng93/BiTA.
| Original language | English |
|---|---|
| Article number | 127305 |
| Journal | Expert Systems with Applications |
| Volume | 279 |
| DOIs | |
| State | Published - 15 Jun 2025 |
| Externally published | Yes |
Keywords
- Draft verification
- LLM
- Lossless acceleration
- Prompt tuning
- Semi-autoregressive generation
- Speculative decoding
Fingerprint
Dive into the research topics of 'BiTA: Bi-directional tuning for lossless acceleration in large language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver