arrow
Return

FlatStor: An Efficient Embedded-Index Based Columnar Data Layout for Multimodal Data Workloads

delete2025-09-01
delete0
PRE
AI
C
Chi Zhang
S
Shihao Zhang
Y
Yunfei Gu
C
Chentao Wu *
J
J. Li
Q
Qin Zhang
X
Xusheng Chen
J
Jie Meng
DOI:10.14778/3772181.3772182delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Modern data lakes have become essential for storing, managing, and analyzing massive amounts of heterogeneous data. As producand multi-purpose access patterns, efficient management of such complexities becomes critical. However, current hybrid storage system-based data lakes face persistent challenges, including synchronization overhead, data correlation disruption, and escalating storage costs due to the involvement of multiple underlying storage systems. While columnar storage, central to data lakes, addresses hybrid-system inefficiencies, it struggles with the complexities of multimodal data storage and multi-purpose access. To tackle these challenges, we analyze access patterns across various scenarios and assess the issues in storing multimodal data. Based on these insights, we propose FlatStor, a FlatBuffers-based columnar Storage format with embedded indexing. It supports point access through indexing and handles multimodal data by vertically partitioning and treating each modality as a byte stream for storage. It also applies FSST compression, reducing storage overhead the access latency by 99.6% and the storage overhead by 91.3% compared to Parquet in inference workloads. Furthermore, FlatStor ing minimal additional overhead.
Keywords:
FlatStor
columnar storage
multimodal data
embedded indexing
data lake

Journal

P
Proceedings of the VLDB Endowment
IF:
3.3
Papers:
556
Citations:
1.2W

Organization

S
shanghai jiao tong university
Scholars:
15.6W
Papers: 11.6W
Citations: 159