TY - GEN
T1 - Efficient string similarity search on disks
AU - Wang, Jinbao
AU - Yang, Donghua
N1 - Publisher Copyright:
© Springer-Verlag Berlin Heidelberg 2015.
PY - 2015
Y1 - 2015
N2 - String similarity search is a basic operation for various applications, such as data cleaning, spell checking, bioinformatics and information integration. Memory based q-gram inverted indexes fail to support string similarity search over large scale string datasets due to the memory limitation, and it can no longer work if the data size grows beyond the memory size. In the era of big data, large string dataset are quite common. Existing external memory method, Behm-Index, only supports length-filter and prefix filter. This paper proposes LPA-Index to reduce I/O cost for better query response time, and LPA-Index is a disk resident index which suffers no limitation on data size compared to memory size. LPA-Index supports multiple filters to reduce query candidates effectively, and it adaptively reads inverted lists during query processing for better I/O performance. Experiment results demonstrate the efficiency of LPA-Index and its advantages over existing state-of-art disk index Behm-Index with regard to I/O cost and query response time.
AB - String similarity search is a basic operation for various applications, such as data cleaning, spell checking, bioinformatics and information integration. Memory based q-gram inverted indexes fail to support string similarity search over large scale string datasets due to the memory limitation, and it can no longer work if the data size grows beyond the memory size. In the era of big data, large string dataset are quite common. Existing external memory method, Behm-Index, only supports length-filter and prefix filter. This paper proposes LPA-Index to reduce I/O cost for better query response time, and LPA-Index is a disk resident index which suffers no limitation on data size compared to memory size. LPA-Index supports multiple filters to reduce query candidates effectively, and it adaptively reads inverted lists during query processing for better I/O performance. Experiment results demonstrate the efficiency of LPA-Index and its advantages over existing state-of-art disk index Behm-Index with regard to I/O cost and query response time.
KW - We would like to encourage you to list your within the abstract section
UR - https://www.scopus.com/pages/publications/84927738247
M3 - 会议稿件
AN - SCOPUS:84927738247
T3 - IFIP Advances in Information and Communication Technology
SP - 48
EP - 55
BT - Intelligent Computation in Big Data Era - International Conference of Young Computer Scientists, Engineers and Educators, ICYCSEE 2015, Proceedings
A2 - Wang, Hongzhi
A2 - Che, Wanxiang
A2 - Qiu, Zhaowen
A2 - Han, Zhongyuan
A2 - Lin, Junyu
A2 - Qi, Haoliang
A2 - Lin, Zeguang
A2 - Kong, Leilei
PB - Springer New York LLC
T2 - International Conference of Young Computer Scientists, Engineers and Educators, ICYCSEE 2015
Y2 - 10 January 2015 through 12 January 2015
ER -