TY - GEN
T1 - Subset Discovery for Entity Matching
AU - Liang, Zheng
AU - Tang, Yafeng
AU - Wang, Hongzhi
AU - Cheng, Haifeng
AU - Ding, Xiaoou
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Entity Matching (EM) is the task of identifying co-referent manifestations from multiple data sources. It has been a long-standing challenge in data integration, knowledge graph, and recommendation systems. In recent years, Machine Learning models are taking off in commercial EM pipelines. However, we observe that most challenging EM tasks have complex structures of local distributions, making ML model perform poorly when trained on the whole dataset. To mitigate this issue, we propose a rule-based subset discovery approach to identify multiple subsets. Our Subset Discovery Rule (SDR) are similar to conventional entity blocking rules, but aim at partitioning training data instead of filtering unmatched entity pair candidates. By training one EM model for each subset, the combined EM prediction can capture local data distributions, yielding accurate Entity Matching results. Naive SDR searching takes O(2n) model retraining, while unsupervised clustering methods perform poorly. Therefore, we propose an O(n2logn) Subset Discovery Rule Searching algorithm. We evaluate SDR on 11 EM benchmarks. The empirical results show that SDR is statistically comparable to SOTA EM methods. Compared with Random Forest (RF)-based EM, the F1-score of SDR-RF is 5.6% higher, and compared with Large Language Model (LLM)-based EM, the F1-score of SDR-LLM is 2.7% higher, showcasing its generalization property. Due to its model-agnostic design, SDR is an extensible plug-in for any machine learning-based EM model.
AB - Entity Matching (EM) is the task of identifying co-referent manifestations from multiple data sources. It has been a long-standing challenge in data integration, knowledge graph, and recommendation systems. In recent years, Machine Learning models are taking off in commercial EM pipelines. However, we observe that most challenging EM tasks have complex structures of local distributions, making ML model perform poorly when trained on the whole dataset. To mitigate this issue, we propose a rule-based subset discovery approach to identify multiple subsets. Our Subset Discovery Rule (SDR) are similar to conventional entity blocking rules, but aim at partitioning training data instead of filtering unmatched entity pair candidates. By training one EM model for each subset, the combined EM prediction can capture local data distributions, yielding accurate Entity Matching results. Naive SDR searching takes O(2n) model retraining, while unsupervised clustering methods perform poorly. Therefore, we propose an O(n2logn) Subset Discovery Rule Searching algorithm. We evaluate SDR on 11 EM benchmarks. The empirical results show that SDR is statistically comparable to SOTA EM methods. Compared with Random Forest (RF)-based EM, the F1-score of SDR-RF is 5.6% higher, and compared with Large Language Model (LLM)-based EM, the F1-score of SDR-LLM is 2.7% higher, showcasing its generalization property. Due to its model-agnostic design, SDR is an extensible plug-in for any machine learning-based EM model.
KW - Entity Matching
KW - Information Integration
KW - Subset Discovery
UR - https://www.scopus.com/pages/publications/105043043550
U2 - 10.1007/978-981-95-3827-0_12
DO - 10.1007/978-981-95-3827-0_12
M3 - 会议稿件
AN - SCOPUS:105043043550
SN - 9789819538263
T3 - Lecture Notes in Computer Science
SP - 182
EP - 197
BT - Database Systems for Advanced Applications - 30th International Conference, DASFAA 2025, Proceedings
A2 - Zhu, Feida
A2 - Lim, Ee-peng
A2 - Yu, Philip S.
A2 - Nadamoto, Akiyo
A2 - Shim, Kyuseok
A2 - Ding, Wei
A2 - Zhang, Bingxue
PB - Springer Science and Business Media Deutschland GmbH
T2 - 30th International Conference on Database Systems for Advanced Applications, DASFAA 2025
Y2 - 26 May 2025 through 29 May 2025
ER -