arrow
Return

SVPDSA: Selective View Perception Data Synthesis With Annotations Using Lightweight Diffusion Network

delete2025-01-01
delete0
delete
OA
AI
S
S Raghavendra
V
Vijayalakshmi
V
Vainidhi
S
S. K. Abhilash
V
Venu Madhav Nookala
P
Parveen Kumar
R
Ramyashree Ramyashree
DOI:10.1109/ACCESS.2025.3588542delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The generation of high-quality annotated image datasets with low computational cost and automated labeling is essential for advancing computer vision systems. However, manual labeling of real images is often labor intensive and expensive. To overcome these challenges, proposed a model named SVPDSA, a generic dataset generation model that incorporates residual and attention block pruning, reduced sampling steps, and network quantization. The model comprises two main training phases: Refined LDM((Latent Diffusion Model) retraining and P-Decoder(Perception decoder) training. A compressed Unet-based diffusion model, pre-trained on the LAION-5B dataset, serves as the foundation for efficient text-to-image synthesis. The model is trained for approximately 50,000 iterations with a learning rate of 0.0001, ensuring lightweight yet effective generation. The proposed model efficiently generates diverse synthetic images with high-quality perception annotations. The proposed approach utilizes a lightweight trained diffusion model and extends text-guided image synthesis to perception data generation, ensuring the quality of the generated datasets while offering a flexible solution for label generation. A decoder module is introduced to expand latent code features and generate labeled annotations for tasks such as semantic segmentation, instance segmentation, and depth estimation. Training the decoder requires fewer than 100 manually labeled images, enabling the creation of an infinitely large annotated dataset. Evaluation on the Cityscapes dataset demonstrates that SVPDSA matches or surpasses existing methods like Mask2Former and DatasetDM in key object classes, including cars, buses, and bicycles. It achieves a mean IoU of 42.7 with ResNet-50 and 41.4 with Swin-B using only 9 real images and 38k synthetic samples, showcasing its efficiency in generating high-quality annotations with minimal real data. Deploying the proposed models on edge devices results in less than a 5-second inference time.This research contributes toward building resource-efficient data generation systems suitable for constrained training environments.
Keywords:
Annotation generation
computer vision
depth estimation
diffusion model
edge deployment
instance segmentation
semantic segmentation
synthetic datasets

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

K
kpit technologies, bengaluru, india
Scholars:
4
Papers: 2
Citations: 0
M
Manipal Institute of Technology
Scholars:
371
Papers: 181
Citations: 0
Cited Papers

Cited Papers

errShare
errSave
Masked-attention Mask Transformer for Universal Image Segmentation
err2022-06-01
err0
errOAAI
errBowen Cheng; Ishan Misra; Alexander G. Schwing; Alexander Kirillov; Rohit Girdhar
errShare
errSave
Toward Extreme Image Compression With Latent Feature Guidance and Diffusion Prior
err
IF0
err2025-01-01
err0
PREAI
errZhiyuan Li; Yanhui Zhou; Hao Wei; Chenyang Ge; Jingwen Jiang
errShare
errSave
High-Resolution Image Synthesis with Latent Diffusion Models
err2022-06-01
err0
errOAAI
errRobin Rombach; Andreas Blattmann; Dominik Lorenz; Patrick Esser; Bjorn Ommer
errShare
errSave
errShare
errSave
errShare
errSave
DatasetGAN: Efficient Labeled Data Factory with Minimal Human Effort
err2021-06-01
err0
errOAAI
errYuxuan Zhang; Huan Ling; Jun Gao; Kangxue Yin; Jean-Francois Lafleche; Adela Barriuso; Antonio Torralba; Sanja Fidler
errShare
errSave
Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models
err2023-06-01
err0
errOAAI
errJiarui Xu; Sifei Liu; Arash Vahdat; Wonmin Byeon; Xiaolong Wang; Shalini De Mello
errShare
errSave
researcher View more