返回
INDREX: In-database relation extraction
DOI:10.1016/j.is.2014.11.006.png)
摘要
En 中文
The management of text data has a long-standing history in the human mankind. A particular common task is extracting relations from text Typically, the user performs this task with two separate systems, a relation extraction system and an SQL-based query engine for analytical tasks. During this iterative analytical workflow, the user must frequently ship data between these systems. Worse, the user must learn to manage both systems. Therefore, end users often desire a single system for both analytical and relation extraction tasks. We propose INDREX, a system that provides a single and comprehensive view of the whole process combining both relation extraction and later exploitation with SQL The system permits a data warehouse style extract-transform-load of generic relations extracted from text documents and can support additional text mining analysis libraries or systems. Once generic relations are loaded, the user can define SQL queries on the extracted relations to discover higher level semantics or to join them with other relational data. For executing this powerful task, our system extends the SQL-based analytical capabilities of a columnar-based massively parallel query processing engine with a broad set of user-defined functions and a data model that supports this task Our white-box approach permits INDREX to benefit from built-in query optimization and indexing techniques of the underlaying query execution engine. Applications that support both text mining and analytical workflows leverage new analytical platforms based on the MapReduce framework and its open source Hadoop implementation. We compare our system against this base line. We measure execution times for common workflows and demonstrate orders of magnitude improvement in execution time using INDREX. (C) 2014 Elsevier Ltd. All rights reserved.
Keyword:
Iterative text mining in a RDBMS
Ad-hoc reports from text data
Information extraction
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.9
论文数:
2.8K
被引数:
1.8K
机构
引用论文
The Influence of Presidential Versus Home State Senatorial Preferences on the Policy Output of Judges on the United States District Courts总统与家乡州参议员偏好对美国地区法院法官政策输出的影响

