arrow
Return

The Web as a parallel corpus

delete2003-09-01
delete284
delete
OA
AI
P
Philip Resnik
N
Noah A. Smith
DOI:10.1162/089120103322711578delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Parallel corpora have become an essential resource for work in multilingual natural language processing. In this article, we report on our work using the STRAND system for mining parallel text on the World Wide Web,first reviewing the original algorithm and results and then presenting a set of significant enhancements. These enhancements include the use of supervised learning based on structural features of documents to improve classification performance, a new content based measure of translational equivalence, and adaptation of the system to take advantage of the Internet Archive for mining parallel text from the Web on a large scale. Finally, the value of these techniques is demonstrated in the construction of a significant parallel corpus for a low-density language pair.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Computational Linguistics cover
Computational Linguistics
IF:
5.3
Papers:
837
Citations:
2.7K

Organization

No organization information available