Skip to main navigation Skip to search Skip to main content

Adversarial Attacks Against World Models: Hallucination-Driven Policy Failure

  • Junjian Zhang
  • , Hao Tan
  • , Ruonan Li
  • , Aiping Li*
  • , Zhaoquan Gu*
  • *Corresponding author for this work
  • National University of Defense Technology
  • National Key Laboratory of Advanced Communication Networks
  • Harbin Institute of Technology
  • Peng Cheng Laboratory

Research output: Contribution to journalArticlepeer-review

Abstract

World models have demonstrated powerful environment modeling capabilities in scenarios such as autonomous driving and robotics, but their adversarial security issues remain underexplored, in particular, adversarial risk analysis of world models. To bridge this gap, we systematically reveal the adversarial risks of world models through two dimensions: fundamental robustness verification and spatio-temporal vulnerability exploration. Specifically, in the fundamental robustness verification, we quantitatively certify the model’s defensive boundaries via white-box attack experiments; in the spatial dimension, by decoupling model component dependencies, we design a spatial-oriented gray-box attack, Latent Space Attack, that misleads policies by perturbing only the perception module, overcoming traditional requirements for full model access; in the temporal dimension, we propose a novel temporally correlated attack method, Temporal Enhancement Attack, incorporating cross-frame adversarial loss. By leveraging the model’s temporal memory and predictive capabilities, this approach amplifies single-frame perturbations to influence future decisions. This work conducts the first verification of the adversarial robustness of world models, thereby pointing out a new direction for adversarial research in reinforcement learning systems.

Original languageEnglish
Article number5484
JournalApplied Sciences (Switzerland)
Volume16
Issue number11
DOIs
StatePublished - Jun 2026
Externally publishedYes

Keywords

  • adversarial attacks
  • white-box attacks
  • world models

Fingerprint

Dive into the research topics of 'Adversarial Attacks Against World Models: Hallucination-Driven Policy Failure'. Together they form a unique fingerprint.

Cite this