arrow
Return

Visual content generation from textual description using improved adversarial network

delete2022-09-15
delete1
PRE
AI
V
Varsha Singh *
U
Uma Shanker Tiwary
DOI:10.1007/s11042-022-13720-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
This paper presents an improved adversarial network for visual content generation from textual description. Synthesizing high-quality images from the textual description is the most challenging problem in Computer vision. Existing methods first generate the initial image sketch and then refine that to fine-grained details at different portions of that image. Mostly available text to image generation methods and approaches nearly reflect the meaning of a given text description. But have not successfully generated details and different parts of the objects. As these methods depend on (1) the initial generated image. If the initial image is not generated correctly, the process fails to generate the fine-grained image with details. (2) According to the image's content, each word has a different level of importance; however, similar text representation is used even for different image contents. Here, an improved Adversarial Network based on hyper-parameter optimization to generate fine-grained images is proposed. Inception Score (IS), t-Distributed Stochastic Neighbor Embedding (TSNE) and R-precision as a metric is used to evaluate and refine the initial image automatically. An attention mechanism is used to pay attention to more valuable words of text description to generate more refined sub-parts of the image. For which an attentional module is used to calculate the matching loss of image-text for generator training. The proposed model has been evaluated on the Caltech-UCSD Birds 200 dataset. Results using Inception score, R-precision, and TSNE matrix shows the model performs favourably against state of the art approaches ATT-GAN (2018) and DM-GAN (2019) improving by 25.72% and 19.37% respectively in terms of Inception score.
Keywords:
Text-to-image generation
Generative Adversarial Networks (GANs)
Fine-grained images
Multi-model problem

Journal

Multimedia Tools and Applications cover
Multimedia Tools and Applications
IF:
3
Papers:
1.9W
Citations:
3.2W

Organization

I
Indian Institute of Information Technology Allahabad
Scholars:
852
Papers: 626
Citations: 833