Abstract
Generating high-quality and diverse human images presents a substantial difficulty within the field of computer vision, especially in developing controllable generative models that can utilize input from various modalities. Such models could enable innovative applications like digital human, fashion design, and content creation. In this study, we introduce Multi2Human, a two-stage image synthesis framework for controllable human image generation with multimodal controls. In the first stage, a novel WAvelet-vqVaE (WAVE) architecture is designed to embed human images using a learnable codebook. The WAVE model enhances the conventional Vector Quantized Variational Autoencoder (VQVAE) by integrating wavelets throughout the encoder, thereby enhancing the quality of image reconstruction and synthesis. In the second stage, a new Multimodal Conditioned Diffusion Model (MCDM) is designed to estimate the underlying prior distribution within the discrete latent space using a discrete diffusion process, thus allowing for human image generation conditioned on multimodal controls. Quantitative and qualitative analysis demonstrates that the proposed method has the ability to create high-quality, lifelike full-body human images while satisfying the specified multimodal controls. Our code is available at https://github.com/gxl-groups/Multi2Human.
| Original language | English |
|---|---|
| Article number | 127682 |
| Journal | Neurocomputing |
| Volume | 587 |
| DOIs | |
| State | Published - 28 Jun 2024 |
| Externally published | Yes |
Keywords
- Controllable image generation
- Diffusion model
- Human image
- Multimodal guidance
Fingerprint
Dive into the research topics of 'Multi2Human: Controllable human image generation with multimodal controls'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver