TY - GEN
T1 - Parallel Skyline Query Processing of Massive Incomplete Activity-Trajectories Data
AU - Belhassena, Amina
AU - Hongzhi, Wang
N1 - Publisher Copyright:
© 2023, The Author(s), under exclusive license to Springer Nature Switzerland AG.
PY - 2023
Y1 - 2023
N2 - The big spatial temporal data captured from technology tools produce massive amount of trajectories data collected from GPS devices. The top-k query was proposed by many researchers, on which they used distance and text parameters for processing. However, the information related to text parameter like activity is always not presented due to some reason like lack internet connection. Furthermore, with massive amount of keyword semantic activity-trajectories, user may enter the wrong activity to find its activity-trajectory. Therefore, it’s hard to return the desirable results based on the exact keyword activity. Our previous work proposed an efficient algorithm to handle the trajectory fuzzy problem based on edit distance and activity weight. However, the algorithm proposed does not work with incomplete Trajectory DataBases (TDBs). Therefore, the present investigation focuses on handling the trajectory skyline problem based on distance and frequent activities in incomplete TDB. To accelerate the query processing, the massive trajectory objects is managed through Distributed Mining Trajectory R-Tree (DMTR-Tree index) based on R-tree indexes and inverted lists. Afterward, an efficient algorithm is developed to handle the query. For a rapid computation, a cluster-computing framework of Apache Spark with MapReduce is used. Theoretical analysis and the experimental results show a well agreement and both attest on the higher efficiency of the proposed algorithm.
AB - The big spatial temporal data captured from technology tools produce massive amount of trajectories data collected from GPS devices. The top-k query was proposed by many researchers, on which they used distance and text parameters for processing. However, the information related to text parameter like activity is always not presented due to some reason like lack internet connection. Furthermore, with massive amount of keyword semantic activity-trajectories, user may enter the wrong activity to find its activity-trajectory. Therefore, it’s hard to return the desirable results based on the exact keyword activity. Our previous work proposed an efficient algorithm to handle the trajectory fuzzy problem based on edit distance and activity weight. However, the algorithm proposed does not work with incomplete Trajectory DataBases (TDBs). Therefore, the present investigation focuses on handling the trajectory skyline problem based on distance and frequent activities in incomplete TDB. To accelerate the query processing, the massive trajectory objects is managed through Distributed Mining Trajectory R-Tree (DMTR-Tree index) based on R-tree indexes and inverted lists. Afterward, an efficient algorithm is developed to handle the query. For a rapid computation, a cluster-computing framework of Apache Spark with MapReduce is used. Theoretical analysis and the experimental results show a well agreement and both attest on the higher efficiency of the proposed algorithm.
KW - Distributed processing
KW - Fuzzy
KW - Incomplete data
KW - Skyline trajectory
KW - Top-k spatial keyword queries
UR - https://www.scopus.com/pages/publications/85144818125
U2 - 10.1007/978-3-031-21595-7_14
DO - 10.1007/978-3-031-21595-7_14
M3 - 会议稿件
AN - SCOPUS:85144818125
SN - 9783031215940
T3 - Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
SP - 193
EP - 206
BT - Model and Data Engineering - 11th International Conference, MEDI 2022, Proceedings
A2 - Fournier-Viger, Philippe
A2 - Hassan, Ahmed
A2 - Bellatreche, Ladjel
PB - Springer Science and Business Media Deutschland GmbH
T2 - 11th International Conference on Model and Data Engineering, MEDI 2022
Y2 - 21 November 2022 through 24 November 2022
ER -