Skip to main navigation Skip to search Skip to main content

Mining parallel corpus from sina microblog

  • School of Computer Science and Technology, Harbin Institute of Technology
  • Heilongjiang Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Finding the parallel corpus as a kind of specific type of information from microblogging sites with millions of users, such as Sina Microblog, is a challenging task. This paper investigates the feasibility of mining such data from the username, the hash tag as well as the user relations by three different methods. The initial experiment is encouraging under the current restriction of limited microblog content access.

Original languageEnglish
Title of host publicationProceedings - 2013 International Conference on Asian Language Processing, IALP 2013
Pages99-102
Number of pages4
DOIs
StatePublished - 2013
Externally publishedYes
Event2013 International Conference on Asian Language Processing, IALP 2013 - Urumqi, Xinjiang, China
Duration: 17 Aug 201319 Aug 2013

Publication series

NameProceedings - 2013 International Conference on Asian Language Processing, IALP 2013

Conference

Conference2013 International Conference on Asian Language Processing, IALP 2013
Country/TerritoryChina
CityUrumqi, Xinjiang
Period17/08/1319/08/13

Keywords

  • Follower
  • Hash tag
  • Parallel Corpus Mining
  • Sina Micoblog
  • Username

Fingerprint

Dive into the research topics of 'Mining parallel corpus from sina microblog'. Together they form a unique fingerprint.

Cite this