Skip to main navigation Skip to search Skip to main content

Domain-Invariant Representation Learning for Generalizable Deepfake Speech Detection

  • Harbin Institute of Technology
  • Chongqing University

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advances in generative AI have significantly improved synthetic speech quality in text-to-speech and voice conversion tasks. This progress has made synthetic speech increasingly indistinguishable from human speech, posing substantial challenges for deepfake speech detection. Current state-of-the-art deepfake speech detection methods leverage pretrained self-supervised learning (SSL) models and achieve strong performance on in-domain datasets. However, these methods still suffer from limited generalization capabilities when facing unseen and cross-domain attacks. To address this, we propose Collaborative Multi-view Spoofing Detection (CMSD) framework that effectively combines complementary spectral and waveform representations within SSL models, together with a unified cross-domain evaluation protocol for assessing generalization. Our framework incorporates two key components: Hierarchical Layer Aggregation (HLA) module and Collaborative Progressive Multi-view Fusion (CPMF) module. Specifically, the HLA module aggregates outputs from multiple hidden layers of the SSL model to construct comprehensive representations with multi-level feature information, while the CPMF module employs a two-stage progressive fusion strategy to integrate spectral and waveform features via bidirectional Mamba architecture, thereby enhancing the learning of robust multi-view representations for spoofing detection. Experiments demonstrate that the proposed method achieves state-of-the-art EER on ASVspoof 2021 LA, while also exhibiting superior cross-domain generalization across diverse unseen domains.

Keywords

  • Deepfake speech detection
  • cross-domain generalization
  • multi-view learning
  • self-spervised learning
  • state space models

Fingerprint

Dive into the research topics of 'Domain-Invariant Representation Learning for Generalizable Deepfake Speech Detection'. Together they form a unique fingerprint.

Cite this