TY - GEN
T1 - Micro tells macro
T2 - 24th ACM Multimedia Conference, MM 2016
AU - Chen, Jingyuan
AU - Song, Xuemeng
AU - Nie, Liqiang
AU - Wang, Xiang
AU - Zhang, Hanwang
AU - Chua, Tat Seng
N1 - Publisher Copyright:
© 2016 ACM.
PY - 2016/10/1
Y1 - 2016/10/1
N2 - Micro-videos, a new form of user generated contents (UGCs), are gaining increasing enthusiasm. Popular microvideos have enormous commercial potential in many ways, such as online marketing and brand tracking. In fact, the popularity prediction of traditional UGCs including tweets, web images, and long videos, has achieved good theoretical underpinnings and great practical success. However, little research has thus far been conducted to predict the popularity of the bite-sized videos. This task is nontrivial due to three reasons: 1) micro-videos are short in duration and of low quality; 2) they can be described by multiple heterogeneous channels, spanning from social, visual, acoustic to textual modalities; and 3) there are no available benchmark dataset and discriminant features that are suitable for this task. Towards this end, we present a transductive multi-modal learning model. The proposed model is designed to find the optimal latent common space, unifying and preserving information from different modalities, whereby micro-videos can be better represented. This latent space can be used to alleviate the information insufficiency problem caused by the brief nature of micro-videos. In addition, we built a benchmark dataset and extracted a rich set of popularity-oriented features to characterize the popular micro-videos. Extensive experiments have demonstrated the effectiveness of the proposed model. As a side contribution, we have released the dataset, codes and parameters to facilitate other researchers.
AB - Micro-videos, a new form of user generated contents (UGCs), are gaining increasing enthusiasm. Popular microvideos have enormous commercial potential in many ways, such as online marketing and brand tracking. In fact, the popularity prediction of traditional UGCs including tweets, web images, and long videos, has achieved good theoretical underpinnings and great practical success. However, little research has thus far been conducted to predict the popularity of the bite-sized videos. This task is nontrivial due to three reasons: 1) micro-videos are short in duration and of low quality; 2) they can be described by multiple heterogeneous channels, spanning from social, visual, acoustic to textual modalities; and 3) there are no available benchmark dataset and discriminant features that are suitable for this task. Towards this end, we present a transductive multi-modal learning model. The proposed model is designed to find the optimal latent common space, unifying and preserving information from different modalities, whereby micro-videos can be better represented. This latent space can be used to alleviate the information insufficiency problem caused by the brief nature of micro-videos. In addition, we built a benchmark dataset and extracted a rich set of popularity-oriented features to characterize the popular micro-videos. Extensive experiments have demonstrated the effectiveness of the proposed model. As a side contribution, we have released the dataset, codes and parameters to facilitate other researchers.
KW - Micro-videos
KW - Multi-view learning
KW - Popularity prediction
UR - https://www.scopus.com/pages/publications/84994608088
U2 - 10.1145/2964284.2964314
DO - 10.1145/2964284.2964314
M3 - 会议稿件
AN - SCOPUS:84994608088
T3 - MM 2016 - Proceedings of the 2016 ACM Multimedia Conference
SP - 898
EP - 907
BT - MM 2016 - Proceedings of the 2016 ACM Multimedia Conference
PB - Association for Computing Machinery, Inc
Y2 - 15 October 2016 through 19 October 2016
ER -