Return
Diff-TST: Diffusion model for one-shot text-image style transfer
DOI:10.1016/j.eswa.2024.125747.png)
Abstract
En 中文
In recent years, there have been significant breakthroughs in style transfer that show impressive results, enabling the automatic creation of high-quality and diverse images. However, these models struggle to generate accurate and realistic text images, especially for languages with complex glyphs. Besides, existing methods often depend on a large amount of multimodally labeled data for supervision, limiting their applicability to languages with extensive labeled resources. In this work, we propose a universal one-shot text-image style transfer framework via the diffusion model, Diff-TST, that only requires one word-level content label and one style reference image for any language. To this end, Diff-TST decomposes the content and style features of text images and generates style transfer images by incorporating content and style guidance into the diffusion process. To further address the problem of text generation inaccuracy via the diffusion model, we propose character-wise encoding to encode the text at the character level and introduce positional encoding and cross- attention to align the input text with the content of the generated image. Extensive qualitative and quantitative experimental results show that our method is applicable in multiple languages (i.e, English, Thai, and Kazakh) as well as outperforms existing methods.
Keywords:
One-shot style transfer
Scene text generation
Diffusion model
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W

