arrow
Return

Disentangling Structure and Appearance in ViT Feature Space

delete2023-11-30
delete0
delete
OA
AI
N
Narek Tumanyan *
O
Omer Bar-Tal
S
Shir Amir
S
Shai Bagon
T
Tali Dekel
DOI:10.1145/3630096delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present a method for semantically transferring the visual appearance of one natural image to another. Specifically, our goal is to generate an image in which objects in a source structure image are painted with the visual appearance of their semantically related objects in a target appearance image. To integrate semantic information into our framework, our key idea is to leverage a pre-trained and fixed Vision Transformer (ViT) model. Specifically, we derive novel disentangled representations of structure and appearance extracted fromdeep ViT features. We then establish an objective function that splices the desired structure and appearance representations, interweaving them together in the space of ViT features. Based on our objective function, we propose two frameworks of semantic appearance transfer - Splice, which works by training a generator on a single and arbitrary pair of structure-appearance images, and SpliceNet, a feed-forward real-time appearance transfer model trained on a dataset of images from a specific domain. Our frameworks do not involve adversarial training, nor do they require any additional input information such as semantic segmentation or correspondences. We demonstrate high-resolution results on a variety of in-the-wild image pairs, under significant variations in the number of objects, pose, and appearance. Code and supplementary material are available in our project page: splice-vit.github.io.
Keywords:
Style transfer
real-time style transfer
feature inversion
vision transformers

Journal

ACM Transactions on Graphics cover
ACM Transactions on Graphics
IF:
9.5
Papers:
4.7K
Citations:
3.6W

Organization

W
Weizmann Institute of Science
Scholars:
1.3W
Papers: 1.1W
Citations: 2.3W