arrow
Return

DocDB: A Database for Unstructured Document Analysis

delete2025-08-01
delete0
PRE
AI
Z
Zequn Li *
Z
Zhong, Yuanhao
柴成亮 cover
柴成亮 (Chengliang Chai)
Z
Zhaoze Sun
Y
Yuan, Ye
C
Cao, Lei
DOI:10.14778/3750601.3750678delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Recent studies have developed LLM-powered data systems that enable database-like analysis of unstructured text documents. While LLMs excel at attribute extraction from documents, their high computational costs and latency make extraction operations the primary performance bottleneck. Existing systems typically adopt traditional relational database query optimization strategies, which prove ineffective in minimizing LLM-related expenses. To fill this gap, we propose DocDB, a prototype system that features a bunch of novel optimization strategies designated to unstructured document analysis. First, we employ a two-level index to reduce LLM extraction costs by selectively retrieving and processing only text segments relevant to target attributes. Second, DocDB employs adaptive execution, generating document-specific plans to minimize LLM extraction frequency based on varying per-document attribute extraction costs. With a real-life scenario, we demonstrate that DocDB allows users to analyze unstructured documents accurately and affordably using SQL-like queries. The corresponding video is available at https://youtu.be/8yDIKOBHIOg.

Journal

P
Proceedings of the VLDB Endowment
IF:
3.3
Papers:
563
Citations:
1.2W

Organization

B
Beijing Institute of Technology
Scholars:
5.2K
Papers: 2.1K
Citations: 6.0W
U
University of Arizona
Scholars:
3.6W
Papers: 3.2W
Citations: 980
Cited Papers

Cited Papers

No cited papers available