Skip to main navigation Skip to search Skip to main content

A Template Independent Approach for Web News and Blog Content Extraction

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The Web has become a large platform for information publishing and consuming. Web news and blog are both representative information sources providing convenient ways to keep informed. In addition to the main content, most web pages also contain navigation panels, advertisements, recommended articles etc. Effectively extracting news and blog content and filtering these noises is necessary and challenging. In this paper we propose a news and blog content extraction approach that is portable to different languages and various domains. Our extensive case studies shows that characters which are not anchor texts but contain stop words are more likely to be the genuine content. Our method first traverses the entire DOM tree and count these valid characters attached to each DOM node. Then we step into the most representative child node based on valid characters recursively. And we finally stop at the main content node with a predefined criterion. To validate the approach, we conduct experiments by using online news and blog files randomly selected from well-known Chinese and English websites. Experimental result shows that our method achieves 96% F1-measure on average and outperforms CETR.

Original languageEnglish
Title of host publicationProceedings - 2016 3rd International Conference on Information Science and Control Engineering, ICISCE 2016
EditorsShaozi Li, Yun Cheng, Ying Dai
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages120-125
Number of pages6
ISBN (Electronic)9781509025350
DOIs
StatePublished - 31 Oct 2016
Externally publishedYes
Event3rd International Conference on Information Science and Control Engineering, ICISCE 2016 - Beijing, China
Duration: 8 Jul 201610 Jul 2016

Publication series

NameProceedings - 2016 3rd International Conference on Information Science and Control Engineering, ICISCE 2016

Conference

Conference3rd International Conference on Information Science and Control Engineering, ICISCE 2016
Country/TerritoryChina
CityBeijing
Period8/07/1610/07/16

Keywords

  • Content extraction
  • Template independent
  • Valid characters

Fingerprint

Dive into the research topics of 'A Template Independent Approach for Web News and Blog Content Extraction'. Together they form a unique fingerprint.

Cite this