arrow
Return

Diff-TST: Diffusion model for one-shot text-image style transfer

delete2025-03-01
delete0
PRE
AI
S
S. W. Pang
X
Xinyuan Chen
H
Hongjian Zhan
B
Bing Yin
Y
Yue Lu *
DOI:10.1016/j.eswa.2024.125747delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In recent years, there have been significant breakthroughs in style transfer that show impressive results, enabling the automatic creation of high-quality and diverse images. However, these models struggle to generate accurate and realistic text images, especially for languages with complex glyphs. Besides, existing methods often depend on a large amount of multimodally labeled data for supervision, limiting their applicability to languages with extensive labeled resources. In this work, we propose a universal one-shot text-image style transfer framework via the diffusion model, Diff-TST, that only requires one word-level content label and one style reference image for any language. To this end, Diff-TST decomposes the content and style features of text images and generates style transfer images by incorporating content and style guidance into the diffusion process. To further address the problem of text generation inaccuracy via the diffusion model, we propose character-wise encoding to encode the text at the character level and introduce positional encoding and cross- attention to align the input text with the content of the generated image. Extensive qualitative and quantitative experimental results show that our method is applicable in multiple languages (i.e, English, Thai, and Kazakh) as well as outperforms existing methods.
Keywords:
One-shot style transfer
Scene text generation
Diffusion model

Journal

Expert Systems with Applications cover
Expert Systems with Applications
IF:
7.5
Papers:
2.9W
Citations:
10.2W

Organization

E
east china normal university
Scholars:
3.0W
Papers: 2.1W
Citations: 25
S
Shanghai Artificial Intelligence Laboratory
Scholars:
470
Papers: 258
Citations: 765