Skip to main navigation Skip to search Skip to main content

Pivot-FSOF: A Full-Stack Optimization Framework for real-time DiT-based TTS on edge devices

  • Chenhe Gong
  • , Huan Guo
  • , Wenjie Zhang
  • , Mingjiang Wang*
  • *Corresponding author for this work
  • Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Compute-intensive models have become the paradigm for high-fidelity generative tasks; however, severe memory bandwidth and I/O bottlenecks hinder their deployment on edge devices. To address this issue, we propose an algorithm-hardware co-design optimization framework, utilizing Text-to-Speech (TTS) as a case study. Algorithmically, we identify a “semantic pivot” during timestep accumulation and design a pivot-aware structured pruning strategy, coupled with sensitivity-aware INT8 quantization and mixed-supervision distillation. Architecturally, we implement targeted edge GPU accelerations, including a fused Rotary Positional Embedding (RoPE) CUDA kernel with vectorized memory access, a pointer-based zero-copy memory handoff mechanism, and CUDA Graph recording to eliminate redundant memory transactions and kernel launch overheads. We integrated this optimization framework on the NVIDIA Jetson AGX Orin and conducted extensive experiments. Experimental results demonstrate that the proposed framework reduces the network layers by 27% while maintaining an objective speaker similarity score of 80.45% under the SIM-T evaluation protocol. Furthermore, the framework achieves a 9.13× end-to-end inference speedup over the ONNX Runtime FP32 baseline and a 2.23× acceleration over the TensorRT FP16 industrial baseline.

Original languageEnglish
Article number103915
JournalJournal of Systems Architecture
Volume179
DOIs
StatePublished - Oct 2026

Keywords

  • Algorithm-hardware co-design
  • CUDA optimization
  • Diffusion Transformer
  • Edge development
  • Memory-bound acceleration

Fingerprint

Dive into the research topics of 'Pivot-FSOF: A Full-Stack Optimization Framework for real-time DiT-based TTS on edge devices'. Together they form a unique fingerprint.

Cite this