TY - GEN
T1 - Hyper-parameter Recommendation for Truth Discovery
AU - Chen, Siying
AU - Ding, Xiaoou
AU - Liang, Zheng
AU - Tang, Yafeng
AU - Wang, Hongzhi
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2025.
PY - 2025
Y1 - 2025
N2 - With the increase of data size, it has become particularly important to find the most trustworthy information among the diverse and contradictory data. The process of discerning claims consistent with the truth from various data sources is commonly referred to as truth discovery. While truth discovery has demonstrated commendable outcomes across diverse applications, most truth discovery algorithms are severely affected by hyper-parameters, thereby limiting their overall performance enhancement potential. Consequently, it is imperative to explore hyper-parameter recommendation for optimizing truth discovery. Initially, we advocate for data augmentation on the input dataset of the truth discovery. Given the limited availability of open-source datasets of truth discovery algorithms, employing data augmentation becomes crucial for enhancing data richness while upholding data quality. Subsequently, we propose a hyper-parameter recommendation method grounded in dataset similarity, model-agnostic meta learning and Bayesian optimization. The proposed method entails a multi-step process. First, a preliminary estimation of the hyper-parameter for the truth discovery algorithm is obtained through meta-learning. This initial estimation serves as input for the Bayesian optimization algorithm, which, in turn, predicts the hyper-parameter values tailored to each dataset. Leveraging similarity measures between datasets, the hyper-parameters for the target dataset are then computed. Following the hyper-parameter recommendation phase, the truth discovery algorithm attains optimal hyper-parameters, resulting in a noteworthy performance average improvement of 18.14% according to extensive experiments.
AB - With the increase of data size, it has become particularly important to find the most trustworthy information among the diverse and contradictory data. The process of discerning claims consistent with the truth from various data sources is commonly referred to as truth discovery. While truth discovery has demonstrated commendable outcomes across diverse applications, most truth discovery algorithms are severely affected by hyper-parameters, thereby limiting their overall performance enhancement potential. Consequently, it is imperative to explore hyper-parameter recommendation for optimizing truth discovery. Initially, we advocate for data augmentation on the input dataset of the truth discovery. Given the limited availability of open-source datasets of truth discovery algorithms, employing data augmentation becomes crucial for enhancing data richness while upholding data quality. Subsequently, we propose a hyper-parameter recommendation method grounded in dataset similarity, model-agnostic meta learning and Bayesian optimization. The proposed method entails a multi-step process. First, a preliminary estimation of the hyper-parameter for the truth discovery algorithm is obtained through meta-learning. This initial estimation serves as input for the Bayesian optimization algorithm, which, in turn, predicts the hyper-parameter values tailored to each dataset. Leveraging similarity measures between datasets, the hyper-parameters for the target dataset are then computed. Following the hyper-parameter recommendation phase, the truth discovery algorithm attains optimal hyper-parameters, resulting in a noteworthy performance average improvement of 18.14% according to extensive experiments.
KW - Bayesian optimization
KW - Meta learning
KW - Truth discovery
UR - https://www.scopus.com/pages/publications/85218443775
U2 - 10.1007/978-981-97-5555-4_18
DO - 10.1007/978-981-97-5555-4_18
M3 - 会议稿件
AN - SCOPUS:85218443775
SN - 9789819755547
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 277
EP - 292
BT - Database Systems for Advanced Applications - 29th International Conference, DASFAA 2024
A2 - Onizuka, Makoto
A2 - Lee, Jae-Gil
A2 - Tong, Yongxin
A2 - Xiao, Chuan
A2 - Ishikawa, Yoshiharu
A2 - Lu, Kejing
A2 - Amer-Yahia, Sihem
A2 - Jagadish, H.V.
PB - Springer Science and Business Media Deutschland GmbH
T2 - 29th International Conference on Database Systems for Advanced Applications, DASFAA 2024
Y2 - 2 July 2024 through 5 July 2024
ER -