Return
Investigating How Large Language Models Process Trajectory Data: Formats and Prompting Strategies for Transportation Mode Detection
Y
H
N
H
DOI:10.1007/s41651-026-00275-2.png)
Abstract
En 中文
The effectiveness of large language models (LLMs) in transportation mode detection (TMD) remains underexplored, creating a significant research gap in understanding how these models process trajectory data. This study uses TMD as a diagnostic setting to examine which trajectory formats and prompting strategies best support LLM-based inference in human mobility, and further analyzes misclassification patterns and the trajectory formats whose reasoning is most susceptible to hallucinations. Using the Geolife dataset, we evaluate pre-trained and fine-tuned LLMs across 14 trajectory formats, categorized into overview information, coordinate-based, and spatial encoding. Meanwhile, two response strategies are compared: direct answer and Chain-of-Thought (CoT) reasoning. The results show that fine-tuning improves accuracy across all formats. The coordinate-based format with timestamps achieves the highest accuracy of 81.7% after fine-tuning using the direct answer strategy. The direct answer strategy outperforms the CoT strategy, reaching an average 42.0% improvement in accuracy after fine-tuning. Spatial misclassification rate maps further show that fine-tuning substantially reduces spatial errors under the direct answer strategy. Additionally, we find systematic confusions between similar modes (bus and car) and frequent CoT hallucinations (e.g., factual inaccuracies) that increase misclassification. These findings highlight the potential of LLMs for TMD in the transport domain through advances in location-based service data and fusion, and emphasize the need for enhanced trajectory formats, improved response strategies, and measures to mitigate hallucinations.
Keywords:
Human mobility
Transportation mode detection
Large language models
Spatiotemporal data
Trajectory analysis
Journal
J
IF:
6.8
Papers:
241
Citations:
887
