arrow
返回

SOEDiff: Efficient Distillation for Small Object Editing

delete2025-11-01
delete1
PRE
AI
Y
Yiming Wu
Q
Qihe Pan
Z
Zhen Zhao *
Z
Zicheng Wang
R
Ronghua Liang
DOI:10.1145/3715915delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this article, we delve into a new task known as Small Object Editing (SOE), which focuses on text-based image inpainting within a constrained, small-sized area. Despite the remarkable success have been achieved by current image inpainting approaches, their application to the SOE task generally results in failure cases such as Object Missing, Text-Image Mismatch, and Distortion. These failures stem from the limited use of small-sized objects in training datasets and the down-sampling operations employed by U-Net models, which hinders accurate generation. To overcome these challenges, we introduce a novel training-based approach, SOEDiff, aimed at enhancing the capability of baseline models like StableDiffusion in editing small-sized objects while minimizing training costs. Specifically, our method involves two key components: SO-LoRA, which efficiently fine-tunes low-rank matrices, and Cross-scale score distillation, which leverages high-resolution predictions from the pre-trained teacher diffusion model. Our method presents significant improvements on the test dataset collected from MSCOCO and OpenImage, validating the effectiveness of our proposed method in SOE. In particular, when comparing SOEDiff with SD-I model on the OpenImage-small-val dataset, we observe a 0.99 improvement in CLIP-Score and a reduction of 2.87 in FID.
Keyword:
Diffusion Model
Small Object Editing
Low-rank Adaptation
Distillation

期刊

ACM Transactions on Multimedia Computing Communications and Applications 封面图
ACM Transactions on Multimedia Computing Communications and Applications
IF:
6
论文数:
2.0K
被引数:
5.4K

机构

U
university of hong kong
学者数:
3.6K
论文数: 1.7K
被引数: 0
Z
Zhejiang University of Technology
学者数:
3.2K
论文数: 1.1K
被引数: 3.0W
U
University of Sydney
学者数:
6.5W
论文数: 6.2W
被引数: 90
J
Jilin University
学者数:
8.7W
论文数: 5.6W
被引数: 8.9K
学者 查看更多机构
引用论文

引用论文

LoMOE: Localized Multi-Object Editing via Multi-Diffusion
err2024-10-28
err0
PREAI
errChakrabarty,Goirik; Chandrasekar,Aditya; Hebbalaguppe,Ramya; AP,Prathosh
err分享
err收藏
All are Worth Words: A ViT Backbone for Diffusion Models
err2023-06-01
err0
errOAAI
errFan Bao; Shen Nie; Kaiwen Xue; Yue Cao; Chongxuan Li; Hang Su; Jun Zhu
err分享
err收藏
Imagic: Text-Based Real Image Editing with Diffusion Models
err2023-06-01
err0
errOAAI
errBahjat Kawar; Shiran Zada; Oran Lang; Omer Tov; Huiwen Chang; Tali Dekel; Inbar Mosseri; Michal Irani
err分享
err收藏
学者 查看更多内容