Abstract
Spiking Transformers offer energy-efficient neuromorphic computing but require (Formula presented) memory during Spatio-Temporal Backpropagation (STBP), where L is network depth and T is simulation timesteps. Reversible architectures can reduce this by reconstructing activations during backpropagation, but face two critical limitations when scaling to hierarchical Transformers: (1) they do not support hierarchical spiking Transformers with dimension rescaling, (2) they are incompatible with integer spiking neurons, leading to gradient instability. To overcome these limitations, we propose MT-RevSNN, the first reversible framework achieving (Formula presented) memory for hierarchical spiking Transformers. Specifically, we introduce Information Fusion Blocks and Dual-Factor Scaling (DFS) to enable reversible dimension rescaling and stabilize integer neuron dynamics. For lightweight implementation, we further propose Low-Rank Compressed Spiking Self-Attention (LRC-SSA). Implemented on the hierarchical QKFormer, the proposed MT-RevSNN significantly reduces training memory by 12.4 × and training time by 1.9 × on ImageNet-1K, while achieving comparable accuracy to its QKFormer counterparts. These results demonstrate the potential of reversible architectures for scalable, memory-efficient training of large-scale spiking Transformers. The code is available in the supplementary material.
| Original language | English |
|---|---|
| Article number | 109370 |
| Journal | Neural Networks |
| Volume | 205 |
| DOIs | |
| State | Published - Jan 2027 |
| Externally published | Yes |
Keywords
- Memory efficiency
- Neuromorphic computing
- Reversible architecture
- Spiking neural network
- Spiking transformer
Fingerprint
Dive into the research topics of 'MT-RevSNN: Memory and time efficient reversible framework for hierarchical spiking transformer'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver