Skip to main navigation Skip to search Skip to main content

Enhancing micro-video understanding by harnessing external sounds

  • Liqiang Nie
  • , Xiang Wang
  • , Jianglong Zhang
  • , Xiangnan He
  • , Hanwang Zhang
  • , Richang Hong
  • , Qi Tian
  • Shandong University
  • National University of Singapore
  • Communication University of China
  • Columbia University
  • Hefei University of Technology
  • University of Texas at San Antonio

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Different from traditional long videos, micro-videos are much shorter and usually recorded at a specific place with mobile devices. To better understand the semantics of a micro-video and facilitate downstream applications, it is crucial to estimate the venue where the micro-video is recorded, for example, in a concert or on a beach. However, according to our statistics over two million micro-videos, only 1.22% of them were labeled with location information. For the remaining large number of micro-videos without location information, we have to rely on their content to estimate their venue categories. This is a highly challenging task, as micro-videos are naturally multi-modal (with textual, visual and, acoustic content), and more importantly, the quality of each modality varies greatly for different micro-videos. In this work, we focus on enhancing the acoustic modality for the venue category estimation task. This is motivated by our finding that although the acoustic signal can well complement the visual and textual signal in reflecting a micro-video's venue, its quality is usually relatively lower. As such, simply integrating acoustic features with visual and textual features only leads to suboptimal results, or even adversely degrades the overall performance (cf. the barrel theory). To address this, we propose to compensate the shortest board - the acoustic modality - via harnessing the external sound knowledge. We develop a deep transfer model which can jointly enhance the concept-level representation of micro-videos and the venue category prediction. To alleviate the sparsity problem of unpopular categories, we further regularize the representation learning of micro-videos of the same venue category. Through extensive experiments on a real-world dataset, we show that our model significantly outperforms the state-of-the-art method [47] in terms of both Micro-F1 and Macro-F1 scores by leveraging the external acoustic knowledge.

Original languageEnglish
Title of host publicationMM 2017 - Proceedings of the 2017 ACM Multimedia Conference
PublisherAssociation for Computing Machinery, Inc
Pages1192-1200
Number of pages9
ISBN (Electronic)9781450349062
DOIs
StatePublished - 23 Oct 2017
Externally publishedYes
Event25th ACM International Conference on Multimedia, MM 2017 - Mountain View, United States
Duration: 23 Oct 201727 Oct 2017

Publication series

NameMM 2017 - Proceedings of the 2017 ACM Multimedia Conference

Conference

Conference25th ACM International Conference on Multimedia, MM 2017
Country/TerritoryUnited States
CityMountain View
Period23/10/1727/10/17

Keywords

  • Deep neural network
  • External sound knowledge
  • Micro-video categorization
  • Representation learning

Fingerprint

Dive into the research topics of 'Enhancing micro-video understanding by harnessing external sounds'. Together they form a unique fingerprint.

Cite this