1
Return

FreeEdit: Mask-Free Reference-Based Image Editing With Multi-Modal Instruction

delete2025-11-24
delete0
PRE
AI
R
Runze He
K
Kai Ma
L
Linjiang Huang
S
Shaofei Huang
J
Jialin Gao
X
Xiaoming Wei
J
Jiao Dai
J
Jizhong Han
S
Si Liu
DOI:10.1109/TPAMI.2025.3636582delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Introducing user-specified visual concepts in image editing is highly practical as these concepts convey the user's intent more precisely than text-based descriptions. We propose FreeEdit, a novel approach for achieving such reference-based image editing, which can accurately reproduce the visual concept from the reference image based on user-friendly language instructions. Our approach leverages the multi-modal instruction encoder to encode language instructions to guide the editing process. This implicit way of locating the editing area eliminates the need for manual editing masks. To enhance the reconstruction of reference details, we introduce the Decoupled Residual Refer-Attention (DRRA) module. This module is designed to integrate fine-grained reference features extracted by a detail extractor into the image editing process in a residual way without interfering with the original self-attention. Given that existing datasets are unsuitable for reference-based image editing tasks, particularly due to the difficulty in constructing image triplets that include a reference image, we curate a high-quality dataset, FreeBench, using a newly developed twice-repainting scheme. FreeBench comprises the images before and after editing, detailed editing instructions, as well as a reference image that maintains the identity of the edited object, encompassing tasks such as object addition, replacement, and deletion. By conducting phased training on FreeBench followed by quality tuning, FreeEdit achieves high-quality zero-shot editing through convenient language instructions. We conduct extensive experiments to evaluate the effectiveness of FreeEdit across multiple task types, demonstrating its superiority over existing methods.
Keywords:
Diffusion models
image editing
instruction-driven editing
reference-based editing

Journal

IEEE Transactions on Pattern Analysis and Machine Intelligence cover
IEEE Transactions on Pattern Analysis and Machine Intelligence
IF:
18.6
Papers:
831
Citations:
9.8W

Organization

B
beihang university
Scholars:
5.2K
Papers: 2.0K
Citations: 21
M
meituan
Scholars:
36
Papers: 17
Citations: 12
C
chinese academy of sciences
Scholars:
54.9W
Papers: 44.5W
Citations: 703
Cited Papers

Cited Papers

Citing Papers

Citing Papers