TY - GEN
T1 - Dependency exploitation
T2 - 26th International Joint Conference on Artificial Intelligence, IJCAI 2017
AU - Zhu, Xinge
AU - Li, Liang
AU - Zhang, Weigang
AU - Rao, Tianrong
AU - Xu, Min
AU - Huang, Qingming
AU - Xu, Dong
PY - 2017
Y1 - 2017
N2 - Visual emotion recognition aims to associate images with appropriate emotions. There are different visual stimuli that can affect human emotion from low-level to high-level, such as color, texture, part, object, etc. However, most existing methods treat different levels of features as independent entity without having effective method for feature fusion. In this paper, we propose a unified CNN-RNN model to predict the emotion based on the fused features from different levels by exploiting the dependency among them. Our proposed architecture leverages convolutional neural network (CNN) with multiple layers to extract different levels of features within a multi-task learning framework, in which two related loss functions are introduced to learn the feature representation. Considering the dependencies within the low-level and high-level features, a bidirectional recurrent neural network (RNN) is proposed to integrate the learned features from different layers in the CNN model. Extensive experiments on both Internet images and art photo datasets demonstrate that our method outperforms the state-of-the-art methods with at least 7% performance improvement.
AB - Visual emotion recognition aims to associate images with appropriate emotions. There are different visual stimuli that can affect human emotion from low-level to high-level, such as color, texture, part, object, etc. However, most existing methods treat different levels of features as independent entity without having effective method for feature fusion. In this paper, we propose a unified CNN-RNN model to predict the emotion based on the fused features from different levels by exploiting the dependency among them. Our proposed architecture leverages convolutional neural network (CNN) with multiple layers to extract different levels of features within a multi-task learning framework, in which two related loss functions are introduced to learn the feature representation. Considering the dependencies within the low-level and high-level features, a bidirectional recurrent neural network (RNN) is proposed to integrate the learned features from different layers in the CNN model. Extensive experiments on both Internet images and art photo datasets demonstrate that our method outperforms the state-of-the-art methods with at least 7% performance improvement.
UR - https://www.scopus.com/pages/publications/85031910934
U2 - 10.24963/ijcai.2017/503
DO - 10.24963/ijcai.2017/503
M3 - 会议稿件
AN - SCOPUS:85031910934
T3 - IJCAI International Joint Conference on Artificial Intelligence
SP - 3595
EP - 3601
BT - 26th International Joint Conference on Artificial Intelligence, IJCAI 2017
A2 - Sierra, Carles
PB - International Joint Conferences on Artificial Intelligence
Y2 - 19 August 2017 through 25 August 2017
ER -