arrow
Return

Zero-Shot Text-Driven Dynamic Neural Radiance Fields Stylization

delete2025-01-01
delete0
PRE
AI
W
Wanlin Liang
H
Hongbin Xu
W
Wanshui Gan
康文雄 (Wenxiong Kang)
DOI:10.1109/TMM.2025.3565983delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Text-driven style transfer for Neural Radiance Fields (NeRFs) is an emerging research topic that leverages text descriptions instead of reference style images to apply style transfer. However, existing methods for stylizing NeRFs predominantly struggle to extend to 4D dynamic scenes, due to NeRFs’ inherent limitation to static environments. Moreover, these current methods require training for each specific text input, which limits them to a single style description and significantly hampers generalizability and applications. In this paper, we introduce a novel approach to zero-shot text-driven 4D style transfer that adopts text inputs into the CLIP’s style space with a canonical feature volume. Specifically, using geometric priors from pre-trained dynamic Neural Radiance Fields, we train a canonical feature volume by rendering feature maps under the supervision of a pre-trained VGG encoder. Then we utilize CLIP’s multi-modal embedding to connect the text descriptions with style images and learn a canonical style transformation matrix in CLIP’s feature space. Experiments show that our method achieves zero-shot text-driven style transfer for dynamic neural radiance fields and maintains good multi-view and cross-time consistency.
Keywords:
Dynamic scenes
novel view synthesis
text-driven stylization

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

T
the university of tokyo
Scholars:
5.0K
Papers: 2.3K
Citations: 1
S
south china university of technology
Scholars:
6.7W
Papers: 5.1W
Citations: 85