Skip to main navigation Skip to search Skip to main content

TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-Commerce

  • Zhaoyu Guo
  • , Zhou Zhao*
  • , Weike Jin
  • , Dazhou Wang
  • , Ruitao Liu
  • , Jun Yu
  • *Corresponding author for this work
  • Zhejiang University
  • Alibaba Group Holding Ltd.
  • Hangzhou Dianzi University

Research output: Contribution to journalArticlepeer-review

Abstract

In e-commerce, product related video is important content to introduce product characteristics and attract consumers. Especially in the recommendation system of e-commerce platform, video highlight detection methods are usually adopted to capture the most attractive clips for showing to consumers, so as to improve the click through rate of products. However, the effect of the current research methods applied to the actual scene is not satisfactory. Compared with other video understanding tasks, video highlight detection is relatively abstract and subjective, and it is difficult to make accurate judgment only by using visual information. Consequently, we put forward multi-modal video highlight detection task, which introduces video related linguistic information as supervised information. And we propose a graph-based commodity-aware model to solve multi-modal video highlight detection in e-commerce scene. Our model consists of multi-modal highlight detection stage and graph-based fine-tuning stage, in which we adopt graph aggregation method to fuse multi-source natural language information and introduce effective visual feature composition method for graph convolution network based highlight detection. Besides, we release the largest e-commerce video highlight detection dataset, TaoHighlight, in which the videos and related data are collected from Taobao e-commerce platform. Our model achieves state-of-art in all separate categories and overall dataset of TaoHighlight, which shows the superiority of our model.

Original languageEnglish
Pages (from-to)2606-2616
Number of pages11
JournalIEEE Transactions on Multimedia
Volume24
DOIs
StatePublished - 2022
Externally publishedYes

Keywords

  • Electronic Commerce
  • Multi-modal learning
  • Video highlight detection

Fingerprint

Dive into the research topics of 'TaoHighlight: Commodity-Aware Multi-Modal Video Highlight Detection in E-Commerce'. Together they form a unique fingerprint.

Cite this