Abstract
Considering that the research on Chinese word segmentation and part-of-speech (POS) tagging for Chinese electronic medical record (CEMR) is currently at a blank stage because of the lack of annotated corpus on CEMR, a complete scheme for data preprocessing to corpus annotation was proposed starting from corpus construction on CEMR so as to obtain a higher annotation consistency, and to build corpus with larger scale and higher quality on CEMR. Furthermore, the statistical lexical differences between CEMR, open-domain corpus and English electronic health record were quantified, and the systematic error analysis was performed on a POS tagging model trained on open-domain corpus. The work lays the foundation for the research on natural language processing (NLP) technologies for CEMR.
| Original language | English |
|---|---|
| Pages (from-to) | 609-615 |
| Number of pages | 7 |
| Journal | Gaojishu Tongxin/Chinese High Technology Letters |
| Volume | 24 |
| Issue number | 6 |
| DOIs | |
| State | Published - 1 Jun 2014 |
| Externally published | Yes |
Keywords
- Annotation consistency
- Chinese electronic medical record (CEMR)
- Error analysis
- Part-of-speech tagging
- Statistical lexical differences
Fingerprint
Dive into the research topics of 'Research on Chinese electronic medical record oriented lexical corpus annotation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver