arrow
Return

Enhancing RAG System Performance Through Semantic Layout Chunking

delete2026-01-01
delete0
PRE
AI
M
Man Qin *
Q
Qiang Sun
T
Tim French
W
Wei Liu
DOI:10.1007/978-981-95-4969-6_3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Retrieval-Augmented Generation (RAG) has proven effective in enhancing large language model (LLM) performance on tasks that require up-to-date, private, or domain-specific knowledge. In RAG systems, documents are segmented into chunks and stored in vector databases ready for retrieval. However, while the chunking strategy plays a critical role in overall system performance, it is often overlooked in RAG system implementation. Common strategies include chunking by character or token count, or using recursive splitting. Although recent research has introduced more advanced approaches, these typically rely either on semantic coherence between sentences or on presentational layout cues to determine chunk boundaries. We propose that semantic layout chunking better preserves both structural integrity and semantic flow, particularly in formal documents that follow logical organizational patterns. Our method integrates semantic labels during chunk storage to enable structure retrieval. We evaluated this approach using the Unstructured Document Analysis (UDA) dataset, which contains PDF documents across multiple domains, comparing it against purely semantic and boundary-aware baselines on retrieval accuracy and question-answering accuracy. The results show that our method achieves superior performance compared to existing approaches, demonstrating the value of combining semantic and structural signals for document chunking in RAG systems.
Keywords:
Chunking
Semantic Layout
Retrieval Augment Generation
Unstructured Document Analysis

Journal

A
AI 2025: ADVANCES IN ARTIFICIAL INTELLIGENCE, PT I
IF:
0
Papers:
32
Citations:
0

Organization

U
university of western australia
Scholars:
2.4K
Papers: 1.2K
Citations: 0