Skip to main navigation Skip to search Skip to main content

Retrieval of degraded Chinese document based on fuzzy coding strategy

  • School of Computer Science and Technology, Harbin Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

For the sake of the low recognition rate for degraded Chinese document, the performance of retrieval is not good if directly based on OCR result. This paper presents a new way to improve the performance of retrieval by fuzzy coding strategy. Lots of character classes with similar shapes are clustered and are indexed by pseudo code. For ease of test, this paper also presents a way to generate ground-truth of imaged document and synthesized degraded document image. A true OCR text collection and two synthesized document image collections are used for performance evaluation, and the result confirms the validation of our method.

Original languageEnglish
Title of host publication2012 International Conference on Systems and Informatics, ICSAI 2012
Pages261-264
Number of pages4
DOIs
StatePublished - 2012
Externally publishedYes
Event2012 International Conference on Systems and Informatics, ICSAI 2012 - Yantai, China
Duration: 19 May 201220 May 2012

Publication series

Name2012 International Conference on Systems and Informatics, ICSAI 2012

Conference

Conference2012 International Conference on Systems and Informatics, ICSAI 2012
Country/TerritoryChina
CityYantai
Period19/05/1220/05/12

Keywords

  • Retrieval of degraded Chinese document
  • Synthesis of degraded document
  • fuzzy coding strategy

Fingerprint

Dive into the research topics of 'Retrieval of degraded Chinese document based on fuzzy coding strategy'. Together they form a unique fingerprint.

Cite this