Skip to main navigation Skip to search Skip to main content

Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation

  • Yunxin Li
  • , Haoyuan Shi
  • , Baotian Hu*
  • , Longyue Wang*
  • , Jiashun Zhu
  • , Jinyi Xu
  • , Zhen Zhao
  • , Min Zhang
  • *Corresponding author for this work
  • Harbin Institute of Technology Shenzhen
  • Alibaba Group Holding Ltd.
  • Jilin University
  • Shanghai Artificial Intelligence Laboratory

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Traditional animation generation methods depend on training generative models with human-labelled data, entailing a sophisticated multi-stage pipeline that demands substantial human effort and incurs high training costs. Due to limited prompting plans, these methods typically produce brief, information-poor, and context-incoherent animations. To overcome these limitations and automate the animation process, we pioneer the introduction of large multimodal models (LMMs) as the core processor to build an autonomous animation-making agent, named Anim-Director. This agent mainly harnesses the advanced understanding and reasoning capabilities of LMMs and generative AI tools to create animated videos from concise narratives or simple instructions. Specifically, it operates in three main stages: Firstly, the Anim-Director generates a coherent storyline from user inputs, followed by a detailed director’s script that encompasses settings of character profiles and interior/exterior descriptions, and context-coherent scene descriptions that include appearing characters, interiors or exteriors, and scene events. Secondly, we employ LMMs with the image generation tool to produce visual images of settings and scenes. These images are designed to maintain visual consistency across different scenes using a visual-language prompting method that combines scene descriptions and images of the appearing character and setting. Thirdly, scene images serve as the foundation for producing animated videos, with LMMs generating prompts to guide this process. The whole process is notably autonomous without manual intervention, as the LMMs interact seamlessly with generative tools to generate prompts, evaluate visual quality, and select the best one to optimize the final output. To assess the effectiveness of our framework, we collect varied short narratives and incorporate various Image/video evaluation metrics including visual consistency and video quality. The experimental results and case studies demonstrate the Anim-Director’s versatility and significant potential to streamline animation creation.

Original languageEnglish
Title of host publicationProceedings - SIGGRAPH Asia 2024 Conference Papers, SA 2024
EditorsStephen N. Spencer
PublisherAssociation for Computing Machinery, Inc
ISBN (Electronic)9798400711312
DOIs
StatePublished - 3 Dec 2024
Externally publishedYes
Event2024 SIGGRAPH Asia 2024 Conference Papers, SA 2024 - Tokyo, Japan
Duration: 3 Dec 20246 Dec 2024

Publication series

NameProceedings - SIGGRAPH Asia 2024 Conference Papers, SA 2024

Conference

Conference2024 SIGGRAPH Asia 2024 Conference Papers, SA 2024
Country/TerritoryJapan
CityTokyo
Period3/12/246/12/24

Keywords

  • Animation Generation
  • Autonomous Agent
  • Image Generation
  • Large Multimodal Models
  • Video

Fingerprint

Dive into the research topics of 'Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation'. Together they form a unique fingerprint.

Cite this