arrow
返回

Average Sparse Attention for Dense Video Captioning From Multiperspective Edge-Computing Cameras

delete2024-12-01
delete1
PRE
AI
L
Ling-Hsuan Huang
C
Ching-Hu Lu *
DOI:10.1109/JSYST.2024.3456864delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In recent years, the artificial intelligence of things (AIoT) has accelerated the development of edge computing. Since existing edge computing for dense video captioning has only explored single-camera decision-making, we propose a lightweight image stitching model that uses a proposed inverted pruned residual model to realize multicamera decision-making to generate more accurate captions. Existing dense video captioning uses an intensive attention mechanism, which readily results in the loss of important information. Thus, our study proposes an average sparse attention mechanism such that the resultant dense video-captioning model is better able to focus on important information and improve the quality of its generated captions. The experiments show that the lightweight video stitching model can reduce model parameters by 13.40% and increase frames per second by 28.96% on an edge platform when compared to the latest studies. Furthermore, a dense video caption network with the average sparse attention mechanism yielded improvements of 22.97% for BLEU3, 35.04% for BLEU4, and 7.51% for METEOR.
Keyword:
Cameras
Computational modeling
Attention mechanisms
Image edge detection
Vectors
Accuracy
Image coding
Detectors
Synthesizers
Streaming media
Dense video captioning
edge computing
image stitching
Internet of Things
lightweight neural networks
sparse attention

期刊

I
IEEE Open Journal of Circuits and Systems
IF:
2.4
论文数:
4.5K
被引数:
387

机构

N
national taiwan university of science & technology
学者数:
8.8K
论文数: 8.7K
被引数: 9
引用论文

引用论文

err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Preferences for Redistribution
err
IF0
err2009-03-01
err0
errOAAI
errAlberto Alesina; Paola Giuliano
err分享
err收藏
学者 查看更多内容