TY - GEN
T1 - New Perspectives in Information Retrieval
T2 - 5th International Conference on Electronic Communication and Artificial Intelligence, ICECAI 2024
AU - Wang, Yingshuo
AU - Zhang, Zhaoxin
AU - Yan, Yijin
AU - Yao, Wangjun
AU - Cheng, Yanan
N1 - Publisher Copyright:
© 2024 IEEE.
PY - 2024
Y1 - 2024
N2 - In the digital age, search engines have become the primary means by which users access online information. Identifying and capturing public service websites through search engines pose significant challenges for the Internet, particularly given the frequent updates to website content and adjustments made to search engine algorithms. Traditional methods relying on fixed, single keywords for website discovery fall short in achieving deep excavation and comprehensive coverage of public service websites. This paper innovatively designs a mining method featuring dynamic updating and expansion of the keyword library. First, it employs search engine APIs and web crawling technologies to swiftly gather and consolidate the latest information from various websites. Next, it conducts statistical analysis of keyword frequencies to iteratively generate a set of relevant keywords. Finally, it dynamically adjusts keyword frequencies based on new search outcomes. Experiments demonstrate that our method attains a coverage rate surpassing 95%, outperforming conventional mining approaches. This study provides a robust and highly adaptable framework for the mining of public service websites, offering researchers more precise and timely information retrieval tools. It has been successfully applied in domains such as film, resulting in the establishment of a rich resource database.
AB - In the digital age, search engines have become the primary means by which users access online information. Identifying and capturing public service websites through search engines pose significant challenges for the Internet, particularly given the frequent updates to website content and adjustments made to search engine algorithms. Traditional methods relying on fixed, single keywords for website discovery fall short in achieving deep excavation and comprehensive coverage of public service websites. This paper innovatively designs a mining method featuring dynamic updating and expansion of the keyword library. First, it employs search engine APIs and web crawling technologies to swiftly gather and consolidate the latest information from various websites. Next, it conducts statistical analysis of keyword frequencies to iteratively generate a set of relevant keywords. Finally, it dynamically adjusts keyword frequencies based on new search outcomes. Experiments demonstrate that our method attains a coverage rate surpassing 95%, outperforming conventional mining approaches. This study provides a robust and highly adaptable framework for the mining of public service websites, offering researchers more precise and timely information retrieval tools. It has been successfully applied in domains such as film, resulting in the establishment of a rich resource database.
KW - Dynamic Keyword Mining
KW - Internet Information Retrieval
KW - Public Service Websites
KW - Web Crawling
UR - https://www.scopus.com/pages/publications/85206105992
U2 - 10.1109/ICECAI62591.2024.10675086
DO - 10.1109/ICECAI62591.2024.10675086
M3 - 会议稿件
AN - SCOPUS:85206105992
T3 - 2024 5th International Conference on Electronic Communication and Artificial Intelligence, ICECAI 2024
SP - 90
EP - 95
BT - 2024 5th International Conference on Electronic Communication and Artificial Intelligence, ICECAI 2024
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 31 May 2024 through 2 June 2024
ER -