Abstract
The raw data utilized in training machine learning models faces a potential threat from membership inference attacks. To mitigate this risk, employing synthetic data instead of real data is proved effective in desensitizing the information. We introduce a novel generative model, combining Variational Autoencoder and Generative Adversarial Network, to enhance privacy protection by generating synthetic data. In our approach, discrete variables are encoded by conditional generators, and sampling training is employed to ensure the distribution of synthetic data closely aligning with the real data. The modification of the model structure prompts a refinement of the loss function. We leverage Wasserstein distance with gradient penalty and SNorm to keep the stability of the model training process. Experimental results demonstrate that the efficacy of our model surpasses existing state-of-the-art models in terms of data utility metrics. Notably, in the face of membership inference attacks, the similarity from the results indicates the difficulty when distinguish the real data from synthetic data. It means our model have highlighting capabilities for the privacy protection.
| Original language | English |
|---|---|
| Article number | 112899 |
| Journal | Knowledge-Based Systems |
| Volume | 309 |
| DOIs | |
| State | Published - 30 Jan 2025 |
| Externally published | Yes |
Keywords
- Generative Adversarial Network
- Membership privacy
- Synthetic data
- Tabular data
- Variational Autoencoder
Fingerprint
Dive into the research topics of 'Synthetic data for enhanced privacy: A VAE-GAN approach against membership inference attacks'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver