Skip to main navigation Skip to search Skip to main content

Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning

  • Xian Zhong
  • , Zipeng Li
  • , Shuqin Chen*
  • , Kui Jiang*
  • , Chen Chen
  • , Mang Ye
  • *Corresponding author for this work
  • Wuhan University of Technology
  • Hubei University of Education
  • Wuhan University
  • University of Central Florida

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Video captioning aims to generate natural language sentences that describe the given video accurately. Existing methods obtain favorable generation by exploring richer visual representations in encode phase or improving the decoding ability. However, the long-tailed problem hinders these attempts at low-frequency tokens, which rarely occur but carry critical semantics, playing a vital role in the detailed generation. In this paper, we introduce a novel Refined Semantic enhancement method towards Frequency Diffusion (RSFD), a captioning model that constantly perceives the linguistic representation of the infrequent tokens. Concretely, a Frequency-Aware Diffusion (FAD) module is proposed to comprehend the semantics of low-frequency tokens to break through generation limitations. In this way, the caption is refined by promoting the absorption of tokens with insufficient occurrence. Based on FAD, we design a Divergent Semantic Supervisor (DSS) module to compensate for the information loss of high-frequency tokens brought by the diffusion process, where the semantics of low-frequency tokens is further emphasized to alleviate the long-tailed problem. Extensive experiments indicate that RSFD outperforms the state-of-the-art methods on two benchmark datasets, i.e., MSR-VTT and MSVD, demonstrate that the enhancement of low-frequency tokens semantics can obtain a competitive generation effect. Code is available at https://github.com/lzp870/RSFD.

Original languageEnglish
Title of host publicationAAAI-23 Technical Tracks 3
EditorsBrian Williams, Yiling Chen, Jennifer Neville
PublisherAAAI press
Pages3724-3732
Number of pages9
ISBN (Electronic)9781577358800
DOIs
StatePublished - 27 Jun 2023
Externally publishedYes
Event37th AAAI Conference on Artificial Intelligence, AAAI 2023 - Washington, United States
Duration: 7 Feb 202314 Feb 2023

Publication series

NameProceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023
Volume37

Conference

Conference37th AAAI Conference on Artificial Intelligence, AAAI 2023
Country/TerritoryUnited States
CityWashington
Period7/02/2314/02/23

Fingerprint

Dive into the research topics of 'Refined Semantic Enhancement towards Frequency Diffusion for Video Captioning'. Together they form a unique fingerprint.

Cite this