arrow
Return

A compression-based algorithm for Chinese word segmentation

delete2000-09-01
delete80
delete
OA
AI
W
William J. Teahan
Y
Yingying Wen
R
Rodger J. McNab
I
Ian H. Witten
DOI:10.1162/089120100561746delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Chinese is written without using spaces or other word delimiters. Although a text may be thought of as a corresponding sequence of words, there is considerable ambiguity in the placement of boundaries. Interpreting a text as a sequence of words is beneficial for some information retrieval and storage tasks: for example, full-text search, word-based compression, and keyphrase extraction. We describe a scheme that infers appropriate positions fou woud boundaries using an adaptive language model that is standard in text compression. It is trained on a corpus of presegmented text, and when applied to new text, interpolates word boundaries so as to maximize the compression obtained. This simple and general method performs well with respect to specialized schemes for Chinese language segmentation.
Keywords:
TEXT SEGMENTATION
MODELS
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Computational Linguistics cover
Computational Linguistics
IF:
5.3
Papers:
837
Citations:
2.7K

Organization

No organization information available