TY - GEN
T1 - A peer-to-peer based passive web crawling system
AU - Chen, Qing Cai
AU - Yang, Xiao Hong
AU - Wang, Xiao Long
PY - 2011
Y1 - 2011
N2 - Though the commercial success of search engines and large scale web page crawlers, the problems of page refresh, new URL discovering, large file downloading, distributed multimedia content feature extracting and indexing etc. are still open. The independent working behavior of each crawler makes it very hard to seek solutions for all these problems under the classical web crawler architecture. To address these problems, this paper proposes an innovative client/server based web crawling system. This system consists of a crawler server and a crawler client which work in the search engine and website end respectively. The crawler server registers itself to the client and joins into a temporary peer-to-peer network to cooperate and share downloaded data with other crawler servers. Different from the classical crawlers, the data downloading procedure is initialized by a client. So for the crawler server, this is a passive web crawling system. The main benefits of this system include the capability of timely management web changes for a crawler, the saving of website bandwidth resources, the capability of downloading large files or multimedia content features, and the capability of protection intellectual properties while indexing and searching the content. Our experiments taken on a simulation system show its efficiency and practicability for the real Internet environments.
AB - Though the commercial success of search engines and large scale web page crawlers, the problems of page refresh, new URL discovering, large file downloading, distributed multimedia content feature extracting and indexing etc. are still open. The independent working behavior of each crawler makes it very hard to seek solutions for all these problems under the classical web crawler architecture. To address these problems, this paper proposes an innovative client/server based web crawling system. This system consists of a crawler server and a crawler client which work in the search engine and website end respectively. The crawler server registers itself to the client and joins into a temporary peer-to-peer network to cooperate and share downloaded data with other crawler servers. Different from the classical crawlers, the data downloading procedure is initialized by a client. So for the crawler server, this is a passive web crawling system. The main benefits of this system include the capability of timely management web changes for a crawler, the saving of website bandwidth resources, the capability of downloading large files or multimedia content features, and the capability of protection intellectual properties while indexing and searching the content. Our experiments taken on a simulation system show its efficiency and practicability for the real Internet environments.
KW - Passive web crawler
KW - peer-to-peer network
KW - search engine
UR - https://www.scopus.com/pages/publications/80155198187
U2 - 10.1109/ICMLC.2011.6016959
DO - 10.1109/ICMLC.2011.6016959
M3 - 会议稿件
AN - SCOPUS:80155198187
SN - 9781457703065
T3 - Proceedings - International Conference on Machine Learning and Cybernetics
SP - 1878
EP - 1883
BT - Proceedings of 2011 International Conference on Machine Learning and Cybernetics, ICMLC 2011
PB - IEEE Computer Society
T2 - 10th International Conference on Machine Learning and Cybernetics, ICMLC 2011
Y2 - 10 July 2011 through 13 July 2011
ER -