Skip to main navigation Skip to search Skip to main content

An information extraction system for heterogeneous Web source

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Information Extraction is the task of identifying information in texts and converting it into a predefined format. In this paper, we build an information integration system which focuses on the information of computer science teachers in Chinese universities. The target of the system is to automatically extract the useful information from heterogeneous sources and re-organize them into structured format. The system includes 4 main modules: web pages retrieval module, web pages' structure classification module, information extraction module and information updating module. We have successfully applied the system to deal with 107 universities in China which shows the effect of the proposed system.

Original languageEnglish
Title of host publication2010 International Conference on Machine Learning and Cybernetics, ICMLC 2010
PublisherIEEE Computer Society
Pages3287-3292
Number of pages6
ISBN (Print)9781424465262
DOIs
StatePublished - 2010
Externally publishedYes

Publication series

Name2010 International Conference on Machine Learning and Cybernetics, ICMLC 2010
Volume6

Keywords

  • Information Extraction
  • Topical crawler
  • Web mining
  • Web page structure classification

Fingerprint

Dive into the research topics of 'An information extraction system for heterogeneous Web source'. Together they form a unique fingerprint.

Cite this