arrow
返回

Toward Real-Time and Efficient Perception Workflows in Software-Defined Vehicles

delete2024-01-01
delete0
PRE
AI
S
Sumaiya
R
Reza Jafarpourmarzouni
Y
Yichen Luo
S
Sidi Lu
Z
Zheng Dong *
DOI:10.1109/JIOT.2024.3492801delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
With the growing demand for software-defined vehicles (SDVs), deep learning-based perception models have become increasingly important in intelligent transportation systems. However, these models face significant challenges in enabling real-time and efficient SDV solutions due to their substantial computational requirements, which are often unavailable in resource-constrained vehicles. As a result, these models typically suffer from low throughput, high latency, and excessive GPU/memory usage, making them impractical for real-time SDV applications. To address these challenges, our research focuses on optimizing model and workflow performance through the integration of pruning and quantization techniques across various computational environments, utilizing frameworks, such as PyTorch, open neural network exchange (ONNX), ONNX Runtime, and TensorRT. We systematically explore and evaluate three distinct pruning methods in combination with multiprecision quantization workflows (FP32, FP16, and INT8) and present the results based on four evaluation metrics: 1) inference throughput; 2) latency; 3) GPU/memory usage; and 4) accuracy. Our designed techniques, including pruning and quantization, along with optimized workflows, can achieve up to 18x faster inference speed and 16.5x higher throughput, while reducing GPU/memory usage by up to 30%, all with minimal impact on accuracy. Our work suggests using the Torch-ONNX-TensorRT workflow quantized with 16-bit floating point precision (FP16) precision and group pruning as the optimal strategy for maximizing inference performance. It demonstrates great potential in optimizing real-time, efficient perception workflows in SDVs, contributing to the enhanced application of deep learning models in resource-constrained environments.
Keyword:
Computational modeling
Quantization (signal)
Accuracy
Throughput
Real-time systems
YOLO
Optimization
Memory management
Deep learning
Training
16-bit floating point precision (FP16)
accuracy
FP32
GPU/memory usage
INT8
latency
pruning
quantization
real-time
software-defined vehicle (SDV)
throughput
workflow

期刊

IEEE Internet of Things Journal 封面图
IEEE Internet of Things Journal
IF:
8.9
论文数:
1.4W
被引数:
7.8W

机构

W
wayne state university
学者数:
2.0W
论文数: 1.6W
被引数: 17
引用论文

引用论文

暂无论文信息