arrow
Return

BigVectorBench: Heterogeneous Data Embedding and Compound Queries are Essential in Evaluating Vector Databases

delete2025-01-01
delete0
PRE
AI
G
Guoxin Kang
Z
Zhongxin Ge
J
Jingpei Hu
X
Xueya Zhang
王蕾 (Lei Wang)
J
Jianfeng Zhan
DOI:10.14778/3718057.3718078delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Vector databases are designed to effectively store, organize, and retrieve high-dimensional vectors, enabling faster and more accurate querying and analysis. This study highlights that the performance of cutting-edge vector databases hinges on their proficiency in managing heterogeneous data embedding and handling compound queries. The former task revolves around converting varied data types into a cohesive vector format, while the latter involves processing multimodal or single-modal queries with precise constraints. The paper advocates for evaluating these dual tasks within an integrated benchmark framework. However, state-of-the-art vector database benchmarks overlook heterogeneous data embedding and compound queries, creating a gap in evaluating vector database performance.
Keywords:
vector databases
heterogeneous data embedding
compound queries
benchmark framework
high-dimensional vectors

Journal

P
Proceedings of the VLDB Endowment
IF:
3.3
Papers:
556
Citations:
1.2W

Organization

No organization information available