Return
Zero-Shot Text-Driven Dynamic Neural Radiance Fields Stylization
DOI:10.1109/TMM.2025.3565983.png)
Abstract
En 中文
Text-driven style transfer for Neural Radiance Fields (NeRFs) is an emerging research topic that leverages text descriptions instead of reference style images to apply style transfer. However, existing methods for stylizing NeRFs predominantly struggle to extend to 4D dynamic scenes, due to NeRFs’ inherent limitation to static environments. Moreover, these current methods require training for each specific text input, which limits them to a single style description and significantly hampers generalizability and applications. In this paper, we introduce a novel approach to zero-shot text-driven 4D style transfer that adopts text inputs into the CLIP’s style space with a canonical feature volume. Specifically, using geometric priors from pre-trained dynamic Neural Radiance Fields, we train a canonical feature volume by rendering feature maps under the supervision of a pre-trained VGG encoder. Then we utilize CLIP’s multi-modal embedding to connect the text descriptions with style images and learn a canonical style transformation matrix in CLIP’s feature space. Experiments show that our method achieves zero-shot text-driven style transfer for dynamic neural radiance fields and maintains good multi-view and cross-time consistency.
Keywords:
Dynamic scenes
novel view synthesis
text-driven stylization
Journal
IF:
9.7
Papers:
4.5K
Citations:
2.4W

