Skip to main navigation Skip to search Skip to main content

A fishing website identification algorithm based on URL text feature and link relation

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Based on the analysis of the uniform resource location (URL) text data of fishing sites and the characteristics of the network topology composed of fishing websites, a fishing site recognition algorithm based on URL text features and link relation (FAUFL) is proposed to improve the accuracy rate of fishing site recognition. The principle of the algorithm is as below: By using URL text features as input, the random forest algorithm is used to generate the fishing site discrimination algorithm based on URL text features. The related web page group is constructed by using the link relation as input, and the related web page algorithm based on the maximum flow cutting is used to generate the fishing website based on the link discriminant algorithm. By taking the above two kinds of discriminant algorithms' results as input, the further evaluation is conducted by using the Bagging algorithm. The test results show that the accuracy rate of the FAUFL is 99.2%, which is 3.9% higher than that of the URL text feature-based algorithm, and 5.0% higher than that of the link-based algorithm.

Original languageEnglish
Pages (from-to)708-717
Number of pages10
JournalGaojishu Tongxin/Chinese High Technology Letters
Volume27
Issue number8
DOIs
StatePublished - 1 Aug 2017
Externally publishedYes

Keywords

  • Fishing website
  • Fusion algorithm
  • Link relation
  • Text feature
  • Uniform resource location (URL)

Fingerprint

Dive into the research topics of 'A fishing website identification algorithm based on URL text feature and link relation'. Together they form a unique fingerprint.

Cite this