Skip to main navigation Skip to search Skip to main content

ACIGS: An automated large-scale crops image generation system based on large visual language multi-modal models

  • Harbin Institute of Technology
  • Harbin Institute of Technology Shenzhen

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Smart agriculture requires an extensive convergence of information technology and agriculture. Attaining intelligence mandates an enormous amount of data to train models. However, it is challenging to acquire a large number of crop image data, limiting the application and growth of computer vision technology in agriculture. To address this problem, we designed a crop image generation system that combines a large language model with visual language multi-modal large models to augment the scale, variety, and resolution of crop image data. First, the system inputs existing real crop images into the visual language multimodal model to extract features and represent crop images in text form. Then, the system passes the crop text representation to the language model for cleaning and processing, which generates prompts to create crop images. The prompts are input into the visual language multi-modal model to generate crop images based on text representation of crops. The resulting crop images undergo image quality evaluation in the visual language multimodal model, and high-quality crop images are saved to the crop image dataset based on the quality evaluation. These steps lead to the formation of the final generated crop image dataset. The experimental results indicate that the crop images generated using the proposed system are similar to but different from the example images. This characteristic enables the expansion of crop data while circumventing redundancy and allowing for resolution control, which is crucial for dense segmentation tasks. Using this method, the existing data can be enlarged up to 7.5 times.

Original languageEnglish
Title of host publication2023 20th Annual IEEE International Conference on Sensing, Communication, and Networking, SECON 2023
PublisherIEEE Computer Society
Pages7-13
Number of pages7
ISBN (Electronic)9798350300529
DOIs
StatePublished - 2023
Event20th Annual IEEE International Conference on Sensing, Communication, and Networking, SECON 2023 - Madrid, Spain
Duration: 11 Sep 202314 Sep 2023

Publication series

NameAnnual IEEE Communications Society Conference on Sensor, Mesh and Ad Hoc Communications and Networks workshops
Volume2023-September
ISSN (Print)2155-5486
ISSN (Electronic)2155-5494

Conference

Conference20th Annual IEEE International Conference on Sensing, Communication, and Networking, SECON 2023
Country/TerritorySpain
CityMadrid
Period11/09/2314/09/23

Keywords

  • Automated Systems
  • Crops Image Generation
  • Large Language Model
  • Large Visual Language Multi-modal Model

Fingerprint

Dive into the research topics of 'ACIGS: An automated large-scale crops image generation system based on large visual language multi-modal models'. Together they form a unique fingerprint.

Cite this