Skip to main navigation Skip to search Skip to main content

Semantics-Preserving Contrastive Representation Learning for Encrypted Traffic Classification

  • Yihao Chen
  • , Shihao Peng
  • , Jinchuan Liu
  • , Zehua Zhang
  • , Daojing He*
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Peng Cheng Laboratory
  • Guangdong University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

As network traffic encryption becomes increasingly prevalent, traditional content-based traffic classification methods face severe challenges. Self-supervised learning offers a new paradigm for learning general representations from unlabeled traffic, yet existing methods still exhibit significant shortcomings. First, data augmentation (DA) strategies are often directly transferred from other domains, which disrupts the inherent temporal logic and protocol structure of network traffic, making it difficult to generate semantically consistent positive and negative sample pairs. Second, existing contrastive learning frameworks are insufficiently adaptive to network dynamics, often treating transport-layer fluctuations as noise to be suppressed, resulting in learned representations that are sensitive to network environmental changes and have limited generalization capabilities. To address these issues, this article proposes a self-supervised learning framework based on the TCP reliable transmission mechanism, aiming to learn semantically preserved general representations from the intrinsic dynamics of traffic. The framework first designs a domain-consistent DA paradigm by algorithmically simulating TCP timeout retransmission and fast retransmission mechanisms to generate traffic variants that preserve the original semantics, thereby constructing high-quality positive sample pairs. On this basis, a general representation learning framework based on momentum contrast is constructed, leveraging dynamic negative sample queue (DQ) and a momentum encoder to perform pretraining on large-scale unlabeled traffic, efficiently learning discriminative features that are resistant to network fluctuations and focused on application-layer semantics. Experiments show that, after fine-tuning, our method outperforms other state-of-the-art network traffic classification methods on multiple encrypted traffic classification benchmark datasets, validating the effectiveness of the proposed augmentation strategy and learning framework.

Original languageEnglish
Pages (from-to)28111-28120
Number of pages10
JournalIEEE Internet of Things Journal
Volume13
Issue number13
DOIs
StatePublished - 1 Jul 2026
Externally publishedYes

Keywords

  • Contrastive learning
  • data augmentation (DA)
  • encrypted traffic classification
  • self-supervised learning

Fingerprint

Dive into the research topics of 'Semantics-Preserving Contrastive Representation Learning for Encrypted Traffic Classification'. Together they form a unique fingerprint.

Cite this