Return
Optimizing CNN-Based Video Summarization With Digital Twin Framework for Adaptive Probability Thresholding
DOI:10.1109/MCOMSTD.2025.3642934.png)
Abstract
En 中文
The rapid growth of video data necessitates efficient summarization techniques to extract critical insights while minimizing computational overhead. This study investigates three pre-trained CNN models-AlexNet (61 M parameters), GoogLeNet (6.8M), and SqueezeNet (1.24M)-with respective inference speeds of 23.8 ms/frame, 18.4 ms/frame, and 12.6 ms/frame. These computational differences significantly affect real-time summarization performance. A probability-driven thresholding mechanism was introduced to optimize object detection, effectively reducing false positives and enhancing classification accuracy. Among the models, GoogLeNet achieved the best trade-off between accuracy (98.96%) and computational efficiency, while AlexNet attained the highest detection accuracy (99.12%) but exhibited a higher false detection rate (5.6%). SqueezeNet, though lightweight and resource-efficient, reached an accuracy of 94.26% but faced limitations in fine-grained recognition tasks. Probability versus correct detection analysis indicated that a threshold of 0.7 maximized accuracy while reducing misclassifications. The proposed adaptive thresholding framework demonstrates strong applicability in real-world scenarios such as wildlife monitoring, intelligent surveillance, and IoT-enabled video analytics. By bridging CNN-based architectures with scalable IoT-driven solutions, this research advances the development of robust, accurate, and efficient video summarization systems for next-generation intelligent environments.
Keywords:
Videos
Computational modeling
Accuracy
Real-time systems
Feature extraction
Computational efficiency
Convolutional neural networks
Deep learning
Adaptation models
Object detection
Journal
I
IF:
0
Papers:
196
Citations:
0

