TY - GEN
T1 - Identifying and Removing the Ghosts of Reproducibility in Service Recommendation Research
AU - Jiang, Tianyu
AU - Liu, Mingyi
AU - Tu, Zhiying
AU - Wang, Zhongjie
N1 - Publisher Copyright:
© 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2023
Y1 - 2023
N2 - A service recommendation system is an information system that helps build mashups quickly to implement new features in response to environmental changes. With the development of deep learning (DL) in recent years, more researchers have started using DL-based methods to solve service recommendation problem and have achieved remarkable results. However, these works have some common deficiencies on non-unified dataset, pre-trained model, evaluation protocol, and experiment environment. These issues will disrupt evaluating the performance of models accurately and make reproducing them difficult. To solve these problems, we propose a service mashup recommendation benchmark (SMRB) that provides a standard environment to enhance comparability between models and credibility of results. We implement eight models (five from top service computing conferences and journals and three created by ourselves) based on SMRB and compare their performance, which proves the effectiveness of SMRB. After analyzing these results, we found that most DL-based models do not perform as well as they promise; instead, the simplest Multilayer perceptron (MLP) models perform better after tuning the parameters, which inspires us to re-examine whether the particular structure of the model can be helpful for the intended purpose and whether it can really improve the performance of the recommendation.
AB - A service recommendation system is an information system that helps build mashups quickly to implement new features in response to environmental changes. With the development of deep learning (DL) in recent years, more researchers have started using DL-based methods to solve service recommendation problem and have achieved remarkable results. However, these works have some common deficiencies on non-unified dataset, pre-trained model, evaluation protocol, and experiment environment. These issues will disrupt evaluating the performance of models accurately and make reproducing them difficult. To solve these problems, we propose a service mashup recommendation benchmark (SMRB) that provides a standard environment to enhance comparability between models and credibility of results. We implement eight models (five from top service computing conferences and journals and three created by ourselves) based on SMRB and compare their performance, which proves the effectiveness of SMRB. After analyzing these results, we found that most DL-based models do not perform as well as they promise; instead, the simplest Multilayer perceptron (MLP) models perform better after tuning the parameters, which inspires us to re-examine whether the particular structure of the model can be helpful for the intended purpose and whether it can really improve the performance of the recommendation.
KW - Benchmark
KW - Reproducibility
KW - Service recommendation
KW - Standard environment
UR - https://www.scopus.com/pages/publications/85163938563
U2 - 10.1007/978-3-031-34560-9_34
DO - 10.1007/978-3-031-34560-9_34
M3 - 会议稿件
AN - SCOPUS:85163938563
SN - 9783031345593
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 577
EP - 593
BT - Advanced Information Systems Engineering - 35th International Conference, CAiSE 2023, Proceedings
A2 - Indulska, Marta
A2 - Reinhartz-Berger, Iris
A2 - Cetina, Carlos
A2 - Pastor, Oscar
PB - Springer Science and Business Media Deutschland GmbH
T2 - 35th International Conference on Advanced Information Systems Engineering, CAiSE 2023
Y2 - 12 June 2023 through 16 June 2023
ER -