Abstract
Recent advances in generative AI have significantly improved synthetic speech quality in text-to-speech and voice conversion tasks. This progress has made synthetic speech increasingly indistinguishable from human speech, posing substantial challenges for deepfake speech detection. Current state-of-the-art deepfake speech detection methods leverage pretrained self-supervised learning (SSL) models and achieve strong performance on in-domain datasets. However, these methods still suffer from limited generalization capabilities when facing unseen and cross-domain attacks. To address this, we propose Collaborative Multi-view Spoofing Detection (CMSD) framework that effectively combines complementary spectral and waveform representations within SSL models, together with a unified cross-domain evaluation protocol for assessing generalization. Our framework incorporates two key components: Hierarchical Layer Aggregation (HLA) module and Collaborative Progressive Multi-view Fusion (CPMF) module. Specifically, the HLA module aggregates outputs from multiple hidden layers of the SSL model to construct comprehensive representations with multi-level feature information, while the CPMF module employs a two-stage progressive fusion strategy to integrate spectral and waveform features via bidirectional Mamba architecture, thereby enhancing the learning of robust multi-view representations for spoofing detection. Experiments demonstrate that the proposed method achieves state-of-the-art EER on ASVspoof 2021 LA, while also exhibiting superior cross-domain generalization across diverse unseen domains.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Deepfake speech detection
- cross-domain generalization
- multi-view learning
- self-spervised learning
- state space models
Fingerprint
Dive into the research topics of 'Domain-Invariant Representation Learning for Generalizable Deepfake Speech Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver