arrow
返回

Using structural contexts to compress semistructured text collections

delete2007-05-01
delete13
delete
OA
AI
G
Gonzalo Navarro
P
Pablo de la Fuente
DOI:10.1016/j.ipm.2006.07.001delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
We describe a compression model for semistructured documents, called Structural Contexts Model (SCM), which takes advantage of the context information usually implicit in the structure of the text. The idea is to use a separate model to compress the text that lies inside each different structure type (e.g., different XML tag). The intuition behind SCM is that the distribution of all the texts that belong to a given structure type should be similar, and different from that of other structure types. We mainly focus on semistatic models, and test our idea using a word-based Huffman method. This is the standard for compressing large natural language text databases, because random access, partial decompression, and direct search of the compressed collection is possible. This variant, dubbed SCM Huff, retains those features and improves Huffman's compression ratios. We consider the possibility that storing separate models may not pay off if the distribution of different structure types is not different enough, and present a heuristic to merge models with the aim of minimizing the total size of the compressed database. This gives an additional improvement over the plain technique. The comparison against existing prototypes shows that, among the methods that permit random access to the collection, SCM Huff achieves the best compression ratios, 2-4% better than the closest alternative. From a purely compression-aimed perspective, we combine SCM with PPM modeling. A separate PPM model is used to compress the text that lies inside each different structure type. The result, SCMPPM, does not permit random access nor direct search in the compressed text, but it gives 2-5% better compression ratios than other techniques for texts longer than 5 MB. (c) 2006 Elsevier Ltd. All rights reserved.
Keyword:
text compression
semistructured documents
compressed text databases
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

I
Information Processing and Management
IF:
6.9
论文数:
5.2K
被引数:
1.4W

机构

暂无机构信息
引用论文

引用论文

MULTILOCUS GENOTYPING OF CRYPTOSPORIDIUM HOMINIS ASSOCIATED WITH DIARRHEA OUTBREAK IN A DAY CARE UNIT IN SÃO PAULO
err2006-04-01
err0
errOAAI
errElenice Messias do Nascimento Gonçalves; Alexandre J. da Silva; Maria Bernadete de Paula Eduardo; Iaiko Horroiva Uemura; Iaci N.S. Moura; Vera L. Pagliusi Castilho; Carlos Eduardo Pereira Corbett
err分享
err收藏
Fast and flexible word searching on compressed text
err2000-04-01
err118
errOAAI
errde Moura, ES; Navarro, G; Ziviani, N; BaezaYates, R
err分享
err收藏
ARITHMETIC CODING FOR DATA-COMPRESSION
err1987-06-01
err1.9K
errOAAI
errWITTEN, IH; NEAL, RM; CLEARY, JG
err分享
err收藏
Measurement of ultraviolet femtosecond pulses using the optical kerr effect
err1992-10-01
err0
PREAI
errH.-St. Albrecht; P. Heist; J. Kleinschmidt; D. van Lap; T. Schröder
err分享
err收藏
Mergers and Acquisitions and Greenfield Foreign Direct Investment in Selected ASEAN Countries
err2019-12-15
err0
errOAAI
errAlireza Tavakol Moghadam; Nur Syazwani Mazlan; Lee Chin; Saifuzzaman Ibrahim
err分享
err收藏
A LOCALLY ADAPTIVE DATA-COMPRESSION SCHEME
err1986-04-01
err334
errOAAI
errBENTLEY, JL; SLEATOR, DD; TARJAN, RE; WEI, VK
err分享
err收藏
学者 查看更多内容