TY - GEN
T1 - Trireme
T2 - 35th ACM Web Conference, WWW 2026
AU - Zhang, Hengtong
AU - Ye, Chen
AU - Wang, Hongzhi
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/4/12
Y1 - 2026/4/12
N2 - Large-scale diffusion models have demonstrated remarkable success across a variety of domains. These models not only exhibit exceptional performance in their primary tasks but also adapt well to downstream applications through the 'pre-train & fine-tune paradigm'. However, the potential misuse of diffusion models for generating unsafe content has raised significant concerns regarding their governance and regulation, necessitating robust unsafe output prevention strategies. Despite the urgent demand for mitigation techniques, a significant challenge persists: once a model is distributed for local deployment or fine-tuning, the model provider and third-party regulators relinquish control over the model's behavior. To address this challenge, we propose a tripartite interactive regulatory scheme (Trireme) to enforce unsafe output prevention for diffusion models. Trireme involves a third-party regulator to monitor the results of the intermediate steps, throughout both the fine-tuning and the sampling phase of the diffusion model. To regulate model behavior, Trireme introduces regulatory signatures and adds additional regulatory layers into vanilla diffusion models. These two components are trained to be tightly coupled with each other, allowing the regulator to control the sampling process of diffusion models with different types of signatures. Through extensive experiments, we demonstrate that the proposed regulatory scheme surpasses existing approaches in terms of preventing the generation of inappropriate content. Moreover, it proves effective even when confronted with multiple adversarial attacks.
AB - Large-scale diffusion models have demonstrated remarkable success across a variety of domains. These models not only exhibit exceptional performance in their primary tasks but also adapt well to downstream applications through the 'pre-train & fine-tune paradigm'. However, the potential misuse of diffusion models for generating unsafe content has raised significant concerns regarding their governance and regulation, necessitating robust unsafe output prevention strategies. Despite the urgent demand for mitigation techniques, a significant challenge persists: once a model is distributed for local deployment or fine-tuning, the model provider and third-party regulators relinquish control over the model's behavior. To address this challenge, we propose a tripartite interactive regulatory scheme (Trireme) to enforce unsafe output prevention for diffusion models. Trireme involves a third-party regulator to monitor the results of the intermediate steps, throughout both the fine-tuning and the sampling phase of the diffusion model. To regulate model behavior, Trireme introduces regulatory signatures and adds additional regulatory layers into vanilla diffusion models. These two components are trained to be tightly coupled with each other, allowing the regulator to control the sampling process of diffusion models with different types of signatures. Through extensive experiments, we demonstrate that the proposed regulatory scheme surpasses existing approaches in terms of preventing the generation of inappropriate content. Moreover, it proves effective even when confronted with multiple adversarial attacks.
KW - abuse prevention
KW - diffusion model
UR - https://www.scopus.com/pages/publications/105038492518
U2 - 10.1145/3774904.3792274
DO - 10.1145/3774904.3792274
M3 - 会议稿件
AN - SCOPUS:105038492518
T3 - WWW 2026 - Proceedings of the ACM Web Conference 2026
SP - 1586
EP - 1594
BT - WWW 2026 - Proceedings of the ACM Web Conference 2026
PB - Association for Computing Machinery, Inc
Y2 - 29 June 2026 through 3 July 2026
ER -