Return
WordCon: Word-Level Typography Control in Visual Text Rendering
DOI:10.1109/tcsvt.2026.3686871.png)
Abstract
En 中文
Visual text rendering represents a fundamental capability of large-scale text-to-image (T2I) models, yet achieving precise word-level controllability remains a significant challenge in this domain. While existing approaches primarily focus on text content accuracy, they often fail to provide fine-grained control over typographic attributes at the word level. To address this limitation, we introduce a comprehensive solution comprising three key components: 1) a novel word-level controlled scene text dataset and benchmark, 2) the Text-Image Alignment (TIA) framework that leverages cross-modal correspondence between textual queries and local image regions through grounding models, and 3) WordCon, a hybrid parameter-efficient fine-tuning (PEFT) method that employs selective parameter reparameterization to enhance both computational efficiency and model portability. The proposed framework incorporates additive supervision mechanisms: a masked loss at the latent level to focus on text regions, and a joint-attention loss at the feature level to promote disentanglement between different words. Extensive experimental evaluations demonstrate that our approach outperforms state-of-the-art methods in both qualitative and quantitative metrics. The proposed method exhibits remarkable versatility, enabling seamless integration with diverse pipelines, including artistic text rendering and image-conditioned text generation. Our datasets, source code, and models will be available for academic research.
Keywords:
Visual text rendering
parameter-efficient fine-tuning
image synthesis
text-image alignment
Journal
IF:
11.1
Papers:
612
Citations:
3.1W

