Skip to main navigation Skip to search Skip to main content

SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions

  • Xiaoyu Liu
  • , Yuxiang Wei
  • , Ming Liu*
  • , Xianhui Lin
  • , Peiran Ren
  • , Xuansong Xie
  • , Wangmeng Zuo
  • *Corresponding author for this work
  • Harbin Institute of Technology
  • Institute for Intelligent Computing
  • Guangdong Artificial Intelligence and Digital Economy Laboratory - Guangzhou

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recent text-to-image generation methods such as ControlNet have achieved remarkable success in controlling image layouts, where the generated images by the default model are constrained to strictly follow the visual conditions (e.g., depth maps). However, in practice, the conditions usually provide only a rough layout, and we argue that the text prompts can more faithfully reflect user intentions. For handling the disagreements between the text prompts and rough visual conditions, we propose a novel text-to-image generation method dubbed SmartControl, which is designed to align well with the text prompts while adaptively keeping useful information from the visual conditions. The key idea of our SmartControl is to relax the constraints on areas that conflict with the text prompts in visual conditions, and two main procedures are required to achieve such a flexible generation. In specific, we extract information from the generative priors of the backbone model (e.g., ControlNet), which effectively represents consistency between the text prompt and visual conditions. Then, a Control Scale Predictor is designed to identify the conflict regions and predict the local control scales. For training the proposed method, a dataset with text prompts and rough visual conditions is constructed. It is worth noting that, even with a limited number (e.g., 1,000–2,000) of training samples, our SmartControl can generalize well to unseen objects. Extensive experiments are conducted on four typical visual condition types, and our SmartControl can achieve a superior performance against state-of-the-art methods. Source code, pre-trained models, and datasets will be publicly available.

Original languageEnglish
Title of host publicationComputer Vision – ECCV 2024 - 18th European Conference, Proceedings
EditorsAleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol
PublisherSpringer Science and Business Media Deutschland GmbH
Pages1-17
Number of pages17
ISBN (Print)9783031731945
DOIs
StatePublished - 2025
Event18th European Conference on Computer Vision, ECCV 2024 - Milan, Italy
Duration: 29 Sep 20244 Oct 2024

Publication series

NameLecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
Volume15106 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th European Conference on Computer Vision, ECCV 2024
Country/TerritoryItaly
CityMilan
Period29/09/244/10/24

Keywords

  • ControlNet
  • Rough Conditions
  • Text-to-Image Generation

Fingerprint

Dive into the research topics of 'SmartControl: Enhancing ControlNet for Handling Rough Visual Conditions'. Together they form a unique fingerprint.

Cite this