arrow
返回

INDREX: In-database relation extraction

delete2015-10-01
delete3
PRE
AI
T
Torsten Kilias *
A
Alexander Löser
P
Periklis Andritsos
DOI:10.1016/j.is.2014.11.006delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
The management of text data has a long-standing history in the human mankind. A particular common task is extracting relations from text Typically, the user performs this task with two separate systems, a relation extraction system and an SQL-based query engine for analytical tasks. During this iterative analytical workflow, the user must frequently ship data between these systems. Worse, the user must learn to manage both systems. Therefore, end users often desire a single system for both analytical and relation extraction tasks. We propose INDREX, a system that provides a single and comprehensive view of the whole process combining both relation extraction and later exploitation with SQL The system permits a data warehouse style extract-transform-load of generic relations extracted from text documents and can support additional text mining analysis libraries or systems. Once generic relations are loaded, the user can define SQL queries on the extracted relations to discover higher level semantics or to join them with other relational data. For executing this powerful task, our system extends the SQL-based analytical capabilities of a columnar-based massively parallel query processing engine with a broad set of user-defined functions and a data model that supports this task Our white-box approach permits INDREX to benefit from built-in query optimization and indexing techniques of the underlaying query execution engine. Applications that support both text mining and analytical workflows leverage new analytical platforms based on the MapReduce framework and its open source Hadoop implementation. We compare our system against this base line. We measure execution times for common workflows and demonstrate orders of magnitude improvement in execution time using INDREX. (C) 2014 Elsevier Ltd. All rights reserved.
Keyword:
Iterative text mining in a RDBMS
Ad-hoc reports from text data
Information extraction
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Enterprise Information Systems 封面图
Enterprise Information Systems
IF:
3.9
论文数:
2.8K
被引数:
1.8K

机构

B
berliner hochschule fur technik
学者数:
132
论文数: 104
被引数: 0
T
Technical University of Berlin
学者数:
1.3W
论文数: 1.1W
被引数: 18
U
University of Lausanne
学者数:
2.5W
论文数: 2.0W
被引数: 3.0W
学者 查看更多机构
引用论文

引用论文

Beyond search: Retrieving complete tuples from a text-database
err2013-01-23
err2
PREAI
errLoeser, Alexander; Nagel, Christoph; Pieper, Stephan; Boden, Christoph
err分享
err收藏
Microscope – A space mission to test the equivalence principle
err2010-01-06
err0
errOAAI
errMeike List; Hanns Selig; Stefanie Bremer; Claus Lämmerzahl
err分享
err收藏
Breaking the Memory Wall in MonetDB
err2008-12-01
err143
errOAAI
errBoncz, Peter A.; Kersten, Martin L.; Manegold, Stefan
err分享
err收藏
Tracking the Temporal Evolution of a Perceptual Judgment Using a Compelled-Response Task
err2011-06-08
err0
errOAAI
errS. Shankar; D. P. Massoglia; D. Zhu; M. G. Costello; T. R. Stanford; E. Salinas
err分享
err收藏
Plaque Growth and Removal With Daily Toothbrushing
err1979-12-01
err0
PREAI
errManuel De la Rosa R.; J. Zacarias Guerra; Dennis A. Johnston; Arthur W. Radike
err分享
err收藏
学者 查看更多内容